如何解决Python中的UnicodeEncodeError及wkhtmltopdf协议未知错误
解决CSV中HTML转PDF的两类错误
错误1:UnicodeEncodeError: 'charmap' codec can't encode character '\u202f'
问题根源
写入HTML文件时未指定编码,Windows系统默认采用cp1252编码,无法处理HTML中的特殊Unicode字符(如\u202f窄空格)。
修复方案
打开HTML文件时显式指定encoding='utf-8',确保特殊字符能正确写入:
with open(html_file_path, 'w', encoding='utf-8') as html_file: html_file.write(html_code)
错误2:OSError: wkhtmltopdf reported an error: Exit with code 1 due to network error: ProtocolUnknownError
问题根源
- 代码中
wkhtmltopdf的路径被换行分割,导致配置路径无效; - HTML内容中包含无法识别的资源协议,触发网络解析错误;
- 写入临时HTML文件再转换的流程,可能引发wkhtmltopdf的路径解析异常。
修复方案
- 修复wkhtmltopdf路径:确保配置路径是完整单行字符串,避免换行:
config = pdfkit.configuration(wkhtmltopdf=r"C:\Program Files\wkhtmltopdf\bin\wkhtmltopdf.exe")
- 禁用外部资源加载:在转换选项中添加禁用外部链接和图片的配置,规避网络相关错误:
options = { 'page-height': '1500.00', 'page-width': '210.00', 'encoding': 'utf8', 'disable-external-links': None, 'disable-images': None }
- 直接用HTML字符串转换:跳过临时HTML文件的写入步骤,使用
pdfkit.from_string直接转换,减少文件IO和路径问题。
修改后的完整代码
import csv import os import pdfkit # 修正原代码中的pdfkitc笔误 # CSV文件路径 csv_file_path = r"C:\Users\jdayao\Desktop\html_to_pdf\emails_body_202303211635.csv" # PDF输出目录 pdf_dir_path = r"C:\Users\jdayao\Desktop\sample_html" # 确保输出目录存在 os.makedirs(pdf_dir_path, exist_ok=True) # 转换配置:添加禁用外部资源选项 options = { 'page-height': '1500.00', 'page-width': '210.00', 'encoding': 'utf8', 'disable-external-links': None, 'disable-images': None } # 修复wkhtmltopdf路径,确保单行完整 config = pdfkit.configuration(wkhtmltopdf=r"C:\Program Files\wkhtmltopdf\bin\wkhtmltopdf.exe") with open(csv_file_path, newline='', encoding="utf8") as csv_file: reader = csv.reader(csv_file) header = next(reader) # 跳过表头 for row in reader: html_code = row[2] # 生成唯一文件名 filename_base = f'{header[2]}_{row[0]}' pdf_file_path = f'{pdf_dir_path}\\{filename_base}.pdf' try: # 直接用HTML字符串转换,无需临时HTML文件 pdfkit.from_string(html_code, pdf_file_path, configuration=config, options=options) print(f'转换成功:{pdf_file_path}') except Exception as e: # 打印具体错误信息,方便排查 print(f'转换失败:{pdf_file_path},错误信息:{str(e)}') print('所有转换任务完成。')
额外注意事项
- 确认
wkhtmltopdf已正确安装,路径与代码配置完全一致; - 若仍有资源加载问题,可尝试添加
'no-stop-slow-scripts': None或'disable-javascript': None到options中; - 原代码中
import pdfkitc为笔误,需改为import pdfkit,否则会触发模块导入错误。
内容的提问来源于stack exchange,提问作者jaypeeee10
相关产品推荐
相关产品推荐

