如何预处理WebP图片以优化Gotenberg生成的PDF文件大小?
解决Gotenberg转换含WebP的HTML为PDF时体积过大的问题
1. 修正HTML图片标签错误
你当前代码里的<img href="img.webp" />是错误写法,图片标签需用src属性指定资源路径,改为:
<img src="img.webp" />
2. 预处理WebP为PDF原生支持的压缩格式
PDF原生支持JPEG、PNG等格式,手动转换并优化比Gotenberg自动转换更可控:
- 无透明通道的照片类图片:转JPEG,调整质量参数(60-80可平衡画质与体积)
- 带透明通道/矢量类图片:转PNG并启用压缩优化
用Python Pillow库实现的优化代码:
from PIL import Image def optimize_webp_for_pdf(input_webp, output_file, img_format="JPEG", quality=70): with Image.open(input_webp) as img: # JPEG不支持透明通道,用白色背景填充 if img_format == "JPEG" and img.mode in ("RGBA", "LA"): bg = Image.new("RGB", img.size, (255, 255, 255)) bg.paste(img, mask=img.split()[-1]) img = bg save_options = {"quality": quality} if img_format == "PNG": save_options["optimize"] = True img.save(output_file, format=img_format, **save_options) # 调用示例:将WebP转为优化后的JPEG optimize_webp_for_pdf("img.webp", "optimized_img.jpg", quality=70)
修改原代码使用优化后的图片:
import io import requests with io.BytesIO() as tmp_index_html: tmp_index_html.write(b""" <html> <head> <title>My img</title> </head> <body> <img src="optimized_img.jpg" /> </body> </html> """) tmp_index_html.seek(0) with open("optimized_img.jpg", "rb") as img_file: response = requests.post( HTML_TO_PDF_URL, files={ "index.html": tmp_index_html, "optimized_img.jpg": img_file, }, timeout=2400, ) with open("result.pdf", "wb") as pdf: pdf.write(response.content)
3. 调整Gotenberg转换参数
通过表单参数启用更优的压缩策略:
response = requests.post( HTML_TO_PDF_URL, files={ "index.html": tmp_index_html, "optimized_img.jpg": img_file, }, data={ "pdf.quality": "75", # 设置PDF输出质量(0-100) "chromium.enableImagesCompression": "true", # 启用Chromium图片压缩 "chromium.printBackground": "true" # 确保图片完整渲染 }, timeout=2400, )
4. 额外优化:控制PDF中图片的显示尺寸
如果不需要原分辨率,在HTML中指定图片的显示尺寸,避免嵌入冗余像素数据:
<!-- 固定显示尺寸 --> <img src="optimized_img.jpg" width="800" height="600" /> <!-- 自适应页面宽度 --> <style> img { max-width: 100%; height: auto; } </style>
内容的提问来源于stack exchange,提问作者Ryabchenko Alexander
相关产品推荐
相关产品推荐

