Streamlit本地环境文件上传至后端卡住问题排查
问题:添加Chroma DB与文本分块后Streamlit文件上传超时失效
原Streamlit应用与LLM对接功能正常,能返回正确响应。但在添加Chroma DB及文本分块、拆分功能后,前端文件上传功能失效:上传文件时卡在“Uploading to backend”环节,即便设置了300秒超时限制,仍会触发超时错误。
Streamlit前端代码
import streamlit as st import requests # Backend URL BACKEND_URL = "http://127.0.0.1:8000" st.title("⚖️ AI Legal Contract Analyzer") # File Upload uploaded_file = st.file_uploader("Upload a legal document (PDF, DOCX, TXT)", type=["pdf", "docx", "txt"]) if uploaded_file is not None: st.info(f"📄 Uploading file: {uploaded_file.name} ({uploaded_file.size} bytes)") # Detect MIME type if uploaded_file.type: mime_type = uploaded_file.type else: # fallback by extension if uploaded_file.name.endswith(".pdf"): mime_type = "application/pdf" elif uploaded_file.name.endswith(".docx"): mime_type = "application/vnd.openxmlformats-officedocument.wordprocessingml.document" else: mime_type = "text/plain" # Show progress progress_bar = st.progress(0) status_text = st.empty() try: status_text.text("🔄 Processing file...") progress_bar.progress(25) # Read file bytes file_bytes = uploaded_file.read() files = {"file": (uploaded_file.name, file_bytes, mime_type)} status_text.text("📤 Uploading to backend...") progress_bar.progress(50) response = requests.post(f"{BACKEND_URL}/parse", files=files, timeout=300) progress_bar.progress(75) if response.status_code == 200: result = response.json() progress_bar.progress(100) status_text.text("✅ Processing complete!") st.success(f"File '{uploaded_file.name}' parsed successfully!") st.info("📊 **Processing Results:**") st.write(f"- **Chunks created:** {result.get('chunks_stored', 'N/A')}") else: progress_bar.progress(0) status_text.text("❌ Processing failed") st.error(f"Failed to parse the file. Status: {response.status_code}") if response.text: st.error(f"Error details: {response.text}") except requests.exceptions.Timeout: progress_bar.progress(0) status_text.text("⏰ Request timed out") st.error("⏰ The file processing took too long.") except requests.exceptions.ConnectionError: progress_bar.progress(0) status_text.text("🔌 Connection error") st.error("🔌 Could not connect to the backend server.") except Exception as e: progress_bar.progress(0) status_text.text("❌ Unexpected error") st.error(f"❌ An unexpected error occurred: {str(e)}") # Question Answering st.subheader("Ask a Question") user_query = st.text_input("Enter your question about the document") if st.button("Get Answer"): if user_query.strip() == "": st.warning("Please enter a question.") else: response = requests.post(f"{BACKEND_URL}/query", json={"query": user_query}) if response.status_code == 200: st.write("### Answer:") st.write(response.json().get("answer")) else: st.error("Error fetching answer from backend.")
报错现象
- 上传进程卡在“Uploading to backend”阶段,无法推进到后续步骤
- 最终触发超时错误提示
排查与解决方案建议
1. 定位后端/parse接口性能瓶颈
- 在后端代码的文本分块、嵌入生成、Chroma DB写入等关键步骤添加日志,记录每个环节的耗时,定位具体慢操作
- 针对大文件优化分块策略:减小单块文本长度、采用并行处理分块,降低整体处理时间
- 改用异步处理+前端轮询模式:后端先接收文件存储,返回任务ID;前端定期调用查询接口获取处理状态,避免长连接超时
2. 优化后端资源与配置
- 若使用本地嵌入模型,优先采用量化版本减小内存占用,或切换为云嵌入API降低本地计算压力
- 将Chroma DB从默认内存模式切换为持久化存储(如SQLite或PostgreSQL),避免内存过载导致写入缓慢
- 调整后端框架(如FastAPI)的超时设置,确保后端处理时间不被框架提前中断
3. 前端请求优化
- 替换同步POST请求为异步流程:先上传文件到后端临时存储,再触发处理任务,前端通过轮询获取结果
- 添加前端重试机制,针对临时网络波动或短时间超时自动重试请求
4. 网络与基础环境检查
- 用
curl直接调用后端/parse接口测试,排除Streamlit前端的请求问题 - 确认后端服务端口未被占用、防火墙未拦截请求,确保前后端通信正常
内容的提问来源于stack exchange,提问作者Mirza Mahad Baig
相关产品推荐
相关产品推荐

