Streamlit实现多PDF文件动态分Tab展示问题求助
动态标签页式PDF查看器解决方案
需求说明
开发Streamlit应用,实现以下功能:
- 支持上传多个PDF文件
- 选中文件时,以文件名为标题动态创建标签页展示PDF内容
- 取消选中文件时,自动关闭对应标签页
原代码问题
原代码未实现动态添加/移除标签页的预期效果,核心问题:
st.tabs使用方式错误,循环绑定内容的逻辑不符合Streamlit组件渲染规则- 未处理文件取消选中的状态同步,也未利用Session State保存上传文件的状态
原代码如下:
import streamlit as st import os from PyPDF2 import PdfReader from io import BytesIO # Function to read PDF and return content def read_pdf(file_path): # Replace with your PDF reading logic pdf = PdfReader(file_path) return pdf # Function to display PDF content def display_pdf(pdf): num_pages = pdf.page_count # Display navigation input for selecting a specific page page_number = st.number_input( "Enter page number", value=1, min_value=1, max_value=num_pages ) # Display an image for the selected page image_bytes = pdf[page_number - 1].get_pixmap().tobytes() st.image(image_bytes, caption=f"Page {page_number} of {num_pages}") # Main Streamlit app code st.title("PDF Viewer App") # Get uploaded files selected_files = st.file_uploader("Upload PDF files", type=["pdf"], accept_multiple_files=True) # List to store content of all pages for each file all_files_content = [] # Dictionary to store selected file content selected_file_content = {} # Iterate over uploaded files for uploaded_file in selected_files: file_content = uploaded_file.read() temp_file_path = f"./temp/{uploaded_file.name}" os.makedirs(os.path.dirname(temp_file_path), exist_ok=True) with open(temp_file_path, "wb") as temp_file: temp_file.write(file_content) if uploaded_file.type == "application/pdf": # Read PDF and store content pdf = read_pdf(temp_file_path) all_files_content.append(pdf) # Display PDF content selected_file_content[uploaded_file.name] = pdf # Cleanup: Remove the temporary file os.remove(temp_file_path) # Create tabs dynamically for each file with st.tabs(list(selected_file_content.keys())): for file_name, pdf in selected_file_content.items(): display_pdf(pdf)
修正后的代码
import streamlit as st from PyPDF2 import PdfReader from io import BytesIO # 初始化Session State,保存已上传的PDF文件 if "uploaded_pdfs" not in st.session_state: st.session_state.uploaded_pdfs = {} # 读取PDF文件(直接使用BytesIO,无需临时文件) def read_pdf(file_bytes): return PdfReader(BytesIO(file_bytes)) # 展示PDF内容 def display_pdf(pdf, tab_key): num_pages = len(pdf.pages) page_number = st.number_input( "输入页码", value=1, min_value=1, max_value=num_pages, key=f"page_num_{tab_key}" ) # 获取页面图片字节流(需确保已安装pdf2image依赖) page = pdf.pages[page_number - 1] pix = page.get_pixmap() st.image(pix.tobytes(), caption=f"第 {page_number} 页 / 共 {num_pages} 页") st.title("PDF 标签页查看器") # 上传文件组件 uploaded_files = st.file_uploader("上传PDF文件", type=["pdf"], accept_multiple_files=True) # 处理上传的文件,存入Session State for file in uploaded_files: if file.name not in st.session_state.uploaded_pdfs: file_bytes = file.read() pdf = read_pdf(file_bytes) st.session_state.uploaded_pdfs[file.name] = pdf # 显示已上传文件的多选框,控制要展示的文件 available_files = list(st.session_state.uploaded_pdfs.keys()) selected_files = st.multiselect("选择要查看的PDF文件", available_files, default=available_files) # 动态创建标签页 if selected_files: tabs = st.tabs(selected_files) for idx, tab in enumerate(tabs): file_name = selected_files[idx] with tab: display_pdf(st.session_state.uploaded_pdfs[file_name], file_name) else: st.info("请上传并选择PDF文件以查看")
关键改进点
- Session State 持久化:用
st.session_state保存已上传的PDF,避免页面刷新后丢失文件状态 - 动态标签页逻辑:通过
st.tabs返回的标签对象逐个绑定对应PDF的展示内容,符合Streamlit组件渲染规则 - 取消选中自动关闭标签:通过
st.multiselect控制选中文件列表,未选中的文件不会生成标签页,实现自动关闭效果 - 移除临时文件:直接用
BytesIO读取上传文件的字节流,无需创建本地临时文件,更高效安全
内容的提问来源于stack exchange,提问作者Mejdi Dallel
相关产品推荐
相关产品推荐

