如何解决Streamlit中Base64编码显示大体积PDF失败问题
问题分析
你的问题核心在于:大体积PDF转Base64后膨胀约33%,超出了浏览器data URI的长度限制(不同浏览器上限通常在2-8MB区间),同时Streamlit的st.markdown和experimental_dialog组件处理大HTML内容时也存在性能或内部长度瓶颈,导致大文件无法加载。
解决方案
方案1:使用Streamlit原生st.pdf组件(推荐)
Streamlit 1.28及以上版本支持直接渲染PDF文件,完全规避Base64编码的体积膨胀问题,对大文件兼容性拉满。修改代码如下:
import streamlit as st import os @st.experimental_dialog("PDF File Viewer", width="large") def show_pdf(file_path): st.pdf(file_path, height=1000) def view_file(file_path): file_extension = os.path.splitext(file_path)[1].lower() if file_extension == '.pdf': show_pdf(file_path) current_directory = os.path.dirname(__file__) document_directory = os.path.join(current_directory, "Document") pdf_file_path = os.path.join(document_directory, "RAG Research.pdf") if st.button("View PDF"): view_file(pdf_file_path)
方案2:通过静态文件服务加载PDF
如果你的Streamlit版本较低,无法使用st.pdf,可以利用Streamlit内置的静态文件服务绕开data URI限制:
- 在项目根目录创建
.streamlit/static文件夹,将PDF文件复制到该目录 - 修改代码如下:
import streamlit as st import os import shutil @st.experimental_dialog("PDF File Viewer", width="large") def show_pdf(file_name): pdf_url = f"/static/{file_name}" st.markdown(f'<iframe src="{pdf_url}" height=1000 width="100%"></iframe>', unsafe_allow_html=True) def view_file(file_path): file_extension = os.path.splitext(file_path)[1].lower() if file_extension == '.pdf': file_name = os.path.basename(file_path) # 确保文件已同步到静态目录 static_dir = os.path.join(os.path.dirname(__file__), ".streamlit", "static") os.makedirs(static_dir, exist_ok=True) target_path = os.path.join(static_dir, file_name) if not os.path.exists(target_path): shutil.copy(file_path, target_path) show_pdf(file_name) current_directory = os.path.dirname(__file__) document_directory = os.path.join(current_directory, "Document") pdf_file_path = os.path.join(document_directory, "RAG Research.pdf") if st.button("View PDF"): view_file(pdf_file_path)
方案3:Base64转Blob URL加载(仅作参考)
如果必须保留Base64编码方式,可以通过JavaScript将Base64数据转为Blob URL,绕开部分data URI限制:
import streamlit as st import os import base64 @st.experimental_dialog("PDF File Viewer", width="large") def show_pdf(base64_pdf): js_code = f""" <script> var base64Data = "{base64_pdf}"; var byteCharacters = atob(base64Data); var byteNumbers = new Array(byteCharacters.length); for (var i = 0; i < byteCharacters.length; i++) {{ byteNumbers[i] = byteCharacters.charCodeAt(i); }} var byteArray = new Uint8Array(byteNumbers); var blob = new Blob([byteArray], {{type: 'application/pdf'}}); var url = URL.createObjectURL(blob); var iframe = document.createElement('iframe'); iframe.src = url; iframe.height = 1000; iframe.width = '100%'; document.body.appendChild(iframe); </script> """ st.components.v1.html(js_code, height=1050) def view_file(file_path): file_extension = os.path.splitext(file_path)[1].lower() if file_extension == '.pdf': with open(file_path, "rb") as f: base64_pdf = base64.b64encode(f.read()).decode('utf-8') show_pdf(base64_pdf) current_directory = os.path.dirname(__file__) document_directory = os.path.join(current_directory, "Document") pdf_file_path = os.path.join(document_directory, "RAG Research.pdf") if st.button("View PDF"): view_file(pdf_file_path)
内容的提问来源于stack exchange,提问作者Alexander
相关产品推荐
相关产品推荐

