PDFKit转换HTML为PDF文件过大且文本无法检索的问题
问题:PDFKit生成的PDF体积过大且文本不可检索(需保留自定义字体解决)
我有一个小于900KB的静态HTML文件(无需联网),用PDFKit生成约100页PDF后,体积达到30-40MB,远超预期——每页仅包含重复文本和4张小图。同时生成的PDF文本无法检索、高亮,每页像一张图片;但用浏览器打印HTML生成的PDF没有这个问题。
我试过把图片缩小到原尺寸的30%,但PDF体积没有变化。需要实现自动化处理,想知道是否遗漏了PDFKit的相关配置?另外我可直接使用HTML字符串而非读取文件,这是否有帮助?也可以更换无授权要求的工具。
当前使用的工具与代码
安装命令
apt-get install wkhtmltopdf -y pip install pdfkit==1.0.0 pip install pypdf2==2.10.5
生成PDF的Python代码
import pdfkit def html_to_pdf(html_path: str, pdf_path: str): pdfkit.from_file( input=html_path, output_path=pdf_path, configuration=pdfkit.configuration(), options={ 'zoom': '0.9588', # 反复测试得到的合适缩放比例 'disable-smart-shrinking': '', 'page-size': 'Letter', 'orientation': 'Landscape', 'margin-top': '0', 'margin-right': '0', 'margin-left': '0', 'margin-bottom': '0', 'encoding': "UTF-8", }) html_to_pdf(".my_html_file.html", "my_pdf_file.pdf")
定位到的核心问题:自定义字体嵌入方式
我使用自定义OTF字体,转成二进制后以base64编码嵌入HTML头部,相关代码如下:
HTML中的字体嵌入代码
<style> @font-face { font-family: "my_custom_font"; src: url(data:font/woff2;base64, asdlfjsads92932super-long-byte-string-here) format("woff2"); font-weight: normal; font-style: normal } </style> <style> @font-face { font-family: "my_custom_font_bold"; src: url(data:font/woff2;base64, asdlfjsads92932super-long-byte-string-here) format("woff2"); font-weight: normal; font-style: normal } </style>
CSS中引用字体
span { font-family: "my_custom_font", Helvetica, sans-serif; } b { font-family: "my_custom_font_bold"; }
HTML主体重复结构
<div class=offer> <div class="offer_banner"> <div class=text_container> <div class="stateroom_text"><span>Deliver to you</span></div> </div> <div class=text_container> <div class=colored_heading> <div class=colored_heading_child><span class=colored_heading_text>My other text</span></div> </div> </div> </div> <div class="offer_top_content"> <div class=text_container> <div class=greeting_text><span>Text 1</span></div> </div> <hr> <div class=text_container> <div class=offer_text><span>text 2</span></div> </div> <hr> <div class=text_container> <div class=redemption_text><span>Text 3</span></div> </div> <div class=text_container> <div><span class="italics_span">Text 4</span></div> </div> </div> <div class=logo_container> <div style="margin:0 20px 0 20px;text-align:center"><img src="data:image/svg+xml,%3Csvg%20xmlns=%27http%3A//www.w3.org/2000/svg%27%20width=%27128%27%20height=%2730%27%3E%3Crect%20fill-opacity=%270%27/%3E%3C/svg%3E" alt="" style="background-blend-mode:normal!important; background-clip:content-box!important; background-position:50% 50%!important; background-color:rgba(0,0,0,0)!important; background-image:var(--sf-img-5)!important; background-size:100% 100%!important; background-origin:content-box!important; background-repeat:no-repeat!important" > </div> </div> </div>
核心需求
在保留自定义字体的前提下,让PDFKit生成的PDF文本可正常检索;同时解决PDF体积过大的问题。
内容的提问来源于stack exchange,提问作者NateH06
相关产品推荐
相关产品推荐

