使用pdfkit转换HTML为PDF遇ProtocolUnknownError错误的排查与解决
问题
我尝试使用pdfkit包将HTML文件转换为PDF,Python代码如下:
location = os.path.join(files_path, f'user_data/temp/{code_generator(8)}') with app.app_context(): html = render_template('email_templates/invoice.html', invoice=invoice) with open(f"{location}.html",'w',encoding = 'utf-8') as f: f.write(html) pdfkit.from_file(f"{location}.html", f"{location}.pdf")
运行后遇到错误:
Traceback (most recent call last): File "/var/www/ifileshifts/run.py", line 3, in <module> from ifileshifts import app, socketio#, manager File "/var/www/ifileshifts/ifileshifts/__init__.py", line 53, in <module> from ifileshifts.main.routes import main File "/var/www/ifileshifts/ifileshifts/main/routes.py", line 6, in <module> from .functions import first_otp_email, duplicate_handler, share_data, convert_size, get_size, is_integer File "/var/www/ifileshifts/ifileshifts/main/functions.py", line 371, in <module> invoice_genarator() File "/var/www/ifileshifts/ifileshifts/main/functions.py", line 353, in invoice_genarator pdfkit.from_file(f"{location}.html", f"{location}.pdf") File "/usr/local/lib/python3.6/dist-packages/pdfkit/api.py", line 51, in from_file return r.to_pdf(output_path) File "/usr/local/lib/python3.6/dist-packages/pdfkit/pdfkit.py", line 201, in to_pdf self.handle_error(exit_code, stderr) File "/usr/local/lib/python3.6/dist-packages/pdfkit/pdfkit.py", line 155, in handle_error raise IOError('wkhtmltopdf reported an error:\n' + stderr) OSError: wkhtmltopdf reported an error: Exit with code 1 due to network error: ProtocolUnknownError
原因分析
- 本地路径协议问题:wkhtmltopdf处理本地文件时,需要明确使用
file://协议前缀,直接传入本地路径会让它误判为网络请求,触发ProtocolUnknownError。 - 外部资源加载失败:如果HTML模板中引用了远程CSS、图片等资源,服务器环境无法访问这些资源时,也会抛出这类网络错误。
解决方法
方法1:给本地路径添加file://前缀
修改pdfkit.from_file的调用,将本地路径转换为file协议格式:
# Linux系统下直接拼接file://前缀 pdfkit.from_file(f"file://{location}.html", f"{location}.pdf") # 若路径包含特殊字符,可先编码 import urllib.parse encoded_path = urllib.parse.quote(f"{location}.html") pdfkit.from_file(f"file://{encoded_path}", f"{location}.pdf")
方法2:直接使用HTML字符串生成PDF
既然已经通过render_template得到了HTML字符串,无需写入本地文件再读取,直接用pdfkit.from_string更高效,还能避免路径问题:
with app.app_context(): html = render_template('email_templates/invoice.html', invoice=invoice) pdfkit.from_string(html, f"{location}.pdf")
方法3:处理HTML中的外部资源
如果模板依赖远程资源,可通过以下方式解决:
- 将远程资源替换为本地文件;
- 给wkhtmltopdf添加允许加载外部资源的参数,同时确保服务器能访问这些资源:
options = { 'enable-local-file-access': None, 'allow': '*', # 允许加载所有外部资源,可按需限制域名 'quiet': '' } pdfkit.from_file(f"{location}.html", f"{location}.pdf", options=options)
内容的提问来源于stack exchange,提问作者Kute Konok
相关产品推荐
相关产品推荐

