You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HTML中经encodeURIComponent编码的Base64格式PDF提取与有效保存

解决嵌入PDF的HTML提取后格式无效问题

问题背景

从包含嵌入PDF的HTML文件中提取Base64编码字符串,用常规Python代码保存为PDF后,Adobe Reader提示格式无效。经排查,原PDF编码时使用了JavaScript的encodeURIComponent函数,常规Base64解码流程无法正确还原内容。

核心差异

原有Python代码直接对提取的Base64字符串解码,但实际编码流程是:PDF二进制 → Base64编码 → encodeURIComponent编码,因此需要先还原URL编码,再进行Base64解码。

修正后的Python实现

使用BeautifulSoup解析HTML提取embed标签的src属性,先做URL解码,再Base64解码后保存为PDF:

import base64
from urllib.parse import unquote
from bs4 import BeautifulSoup
from pathlib import Path

def extract_and_save_pdf(html_path, output_path):
    # 读取HTML文件
    with open(html_path, 'r', encoding='utf-8') as f:
        html_content = f.read()
    
    # 解析HTML提取embed的src属性
    soup = BeautifulSoup(html_content, 'html.parser')
    embed_tag = soup.find('embed', type='application/pdf')
    if not embed_tag or 'src' not in embed_tag.attrs:
        raise ValueError("未找到有效的PDF嵌入标签")
    
    # 处理src内容:移除前缀、URL解码、Base64解码
    src_value = embed_tag['src']
    b64_encoded = src_value.replace('data:application/pdf;base64,', '')
    # 对应JS的decodeURIComponent进行URL解码
    decoded_url = unquote(b64_encoded)
    # 执行Base64解码
    pdf_content = base64.b64decode(decoded_url)
    
    # 保存为PDF文件
    with open(output_path, 'wb') as f:
        f.write(pdf_content)

if __name__ == "__main__":
    html_file = Path("pdf-file.html")
    output_file = Path(Path.home(), 'Downloads', 'mytest.pdf')
    extract_and_save_pdf(html_file, output_file)

代码说明

  • unquote函数对应JavaScript的decodeURIComponent,还原被URL编码的特殊字符(如%2B还原为+)
  • 使用BeautifulSoup替代手动解析HTML,处理逻辑更健壮
  • 增加基础错误判断,避免因标签缺失导致的运行异常

内容的提问来源于stack exchange,提问作者sfgroups

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 15:25:14