You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Streamlit中Base64编码显示大体积PDF失败问题

问题分析

你的问题核心在于:大体积PDF转Base64后膨胀约33%,超出了浏览器data URI的长度限制(不同浏览器上限通常在2-8MB区间),同时Streamlit的st.markdown和experimental_dialog组件处理大HTML内容时也存在性能或内部长度瓶颈,导致大文件无法加载。


解决方案

方案1:使用Streamlit原生st.pdf组件(推荐)

Streamlit 1.28及以上版本支持直接渲染PDF文件,完全规避Base64编码的体积膨胀问题,对大文件兼容性拉满。修改代码如下:

import streamlit as st
import os

@st.experimental_dialog("PDF File Viewer", width="large")
def show_pdf(file_path):
    st.pdf(file_path, height=1000)

def view_file(file_path):
    file_extension = os.path.splitext(file_path)[1].lower()
    if file_extension == '.pdf':
        show_pdf(file_path)

current_directory = os.path.dirname(__file__)
document_directory = os.path.join(current_directory, "Document")
pdf_file_path = os.path.join(document_directory, "RAG Research.pdf")

if st.button("View PDF"):
    view_file(pdf_file_path)

方案2:通过静态文件服务加载PDF

如果你的Streamlit版本较低,无法使用st.pdf,可以利用Streamlit内置的静态文件服务绕开data URI限制:

  1. 在项目根目录创建.streamlit/static文件夹,将PDF文件复制到该目录
  2. 修改代码如下:
import streamlit as st
import os
import shutil

@st.experimental_dialog("PDF File Viewer", width="large")
def show_pdf(file_name):
    pdf_url = f"/static/{file_name}"
    st.markdown(f'<iframe src="{pdf_url}" height=1000 width="100%"></iframe>', unsafe_allow_html=True)

def view_file(file_path):
    file_extension = os.path.splitext(file_path)[1].lower()
    if file_extension == '.pdf':
        file_name = os.path.basename(file_path)
        # 确保文件已同步到静态目录
        static_dir = os.path.join(os.path.dirname(__file__), ".streamlit", "static")
        os.makedirs(static_dir, exist_ok=True)
        target_path = os.path.join(static_dir, file_name)
        if not os.path.exists(target_path):
            shutil.copy(file_path, target_path)
        show_pdf(file_name)

current_directory = os.path.dirname(__file__)
document_directory = os.path.join(current_directory, "Document")
pdf_file_path = os.path.join(document_directory, "RAG Research.pdf")

if st.button("View PDF"):
    view_file(pdf_file_path)

方案3:Base64转Blob URL加载(仅作参考)

如果必须保留Base64编码方式,可以通过JavaScript将Base64数据转为Blob URL,绕开部分data URI限制:

import streamlit as st
import os
import base64

@st.experimental_dialog("PDF File Viewer", width="large")
def show_pdf(base64_pdf):
    js_code = f"""
    <script>
    var base64Data = "{base64_pdf}";
    var byteCharacters = atob(base64Data);
    var byteNumbers = new Array(byteCharacters.length);
    for (var i = 0; i < byteCharacters.length; i++) {{
        byteNumbers[i] = byteCharacters.charCodeAt(i);
    }}
    var byteArray = new Uint8Array(byteNumbers);
    var blob = new Blob([byteArray], {{type: 'application/pdf'}});
    var url = URL.createObjectURL(blob);
    var iframe = document.createElement('iframe');
    iframe.src = url;
    iframe.height = 1000;
    iframe.width = '100%';
    document.body.appendChild(iframe);
    </script>
    """
    st.components.v1.html(js_code, height=1050)

def view_file(file_path):
    file_extension = os.path.splitext(file_path)[1].lower()
    if file_extension == '.pdf':
        with open(file_path, "rb") as f:
            base64_pdf = base64.b64encode(f.read()).decode('utf-8')
        show_pdf(base64_pdf)

current_directory = os.path.dirname(__file__)
document_directory = os.path.join(current_directory, "Document")
pdf_file_path = os.path.join(document_directory, "RAG Research.pdf")

if st.button("View PDF"):
    view_file(pdf_file_path)

内容的提问来源于stack exchange,提问作者Alexander

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 17:55:09