HTML中经encodeURIComponent编码的Base64格式PDF提取与有效保存
解决嵌入PDF的HTML提取后格式无效问题
问题背景
从包含嵌入PDF的HTML文件中提取Base64编码字符串,用常规Python代码保存为PDF后,Adobe Reader提示格式无效。经排查,原PDF编码时使用了JavaScript的encodeURIComponent函数,常规Base64解码流程无法正确还原内容。
核心差异
原有Python代码直接对提取的Base64字符串解码,但实际编码流程是:PDF二进制 → Base64编码 → encodeURIComponent编码,因此需要先还原URL编码,再进行Base64解码。
修正后的Python实现
使用BeautifulSoup解析HTML提取embed标签的src属性,先做URL解码,再Base64解码后保存为PDF:
import base64 from urllib.parse import unquote from bs4 import BeautifulSoup from pathlib import Path def extract_and_save_pdf(html_path, output_path): # 读取HTML文件 with open(html_path, 'r', encoding='utf-8') as f: html_content = f.read() # 解析HTML提取embed的src属性 soup = BeautifulSoup(html_content, 'html.parser') embed_tag = soup.find('embed', type='application/pdf') if not embed_tag or 'src' not in embed_tag.attrs: raise ValueError("未找到有效的PDF嵌入标签") # 处理src内容:移除前缀、URL解码、Base64解码 src_value = embed_tag['src'] b64_encoded = src_value.replace('data:application/pdf;base64,', '') # 对应JS的decodeURIComponent进行URL解码 decoded_url = unquote(b64_encoded) # 执行Base64解码 pdf_content = base64.b64decode(decoded_url) # 保存为PDF文件 with open(output_path, 'wb') as f: f.write(pdf_content) if __name__ == "__main__": html_file = Path("pdf-file.html") output_file = Path(Path.home(), 'Downloads', 'mytest.pdf') extract_and_save_pdf(html_file, output_file)
代码说明
unquote函数对应JavaScript的decodeURIComponent,还原被URL编码的特殊字符(如%2B还原为+)- 使用
BeautifulSoup替代手动解析HTML,处理逻辑更健壮 - 增加基础错误判断,避免因标签缺失导致的运行异常
内容的提问来源于stack exchange,提问作者sfgroups
相关产品推荐
相关产品推荐

