You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Python中的UnicodeEncodeError及wkhtmltopdf协议未知错误

解决CSV中HTML转PDF的两类错误

错误1:UnicodeEncodeError: 'charmap' codec can't encode character '\u202f'

问题根源

写入HTML文件时未指定编码,Windows系统默认采用cp1252编码,无法处理HTML中的特殊Unicode字符(如\u202f窄空格)。

修复方案

打开HTML文件时显式指定encoding='utf-8',确保特殊字符能正确写入:

with open(html_file_path, 'w', encoding='utf-8') as html_file:
    html_file.write(html_code)

错误2:OSError: wkhtmltopdf reported an error: Exit with code 1 due to network error: ProtocolUnknownError

问题根源

  1. 代码中wkhtmltopdf的路径被换行分割,导致配置路径无效;
  2. HTML内容中包含无法识别的资源协议,触发网络解析错误;
  3. 写入临时HTML文件再转换的流程,可能引发wkhtmltopdf的路径解析异常。

修复方案

  1. 修复wkhtmltopdf路径:确保配置路径是完整单行字符串,避免换行:
config = pdfkit.configuration(wkhtmltopdf=r"C:\Program Files\wkhtmltopdf\bin\wkhtmltopdf.exe")
  1. 禁用外部资源加载:在转换选项中添加禁用外部链接和图片的配置,规避网络相关错误:
options = {
    'page-height': '1500.00',
    'page-width': '210.00',
    'encoding': 'utf8',
    'disable-external-links': None,
    'disable-images': None
}
  1. 直接用HTML字符串转换:跳过临时HTML文件的写入步骤,使用pdfkit.from_string直接转换,减少文件IO和路径问题。

修改后的完整代码

import csv
import os
import pdfkit  # 修正原代码中的pdfkitc笔误

# CSV文件路径
csv_file_path = r"C:\Users\jdayao\Desktop\html_to_pdf\emails_body_202303211635.csv"    
# PDF输出目录
pdf_dir_path = r"C:\Users\jdayao\Desktop\sample_html"    

# 确保输出目录存在
os.makedirs(pdf_dir_path, exist_ok=True)

# 转换配置:添加禁用外部资源选项
options = {
    'page-height': '1500.00',
    'page-width': '210.00',
    'encoding': 'utf8',
    'disable-external-links': None,
    'disable-images': None
}  

# 修复wkhtmltopdf路径,确保单行完整
config = pdfkit.configuration(wkhtmltopdf=r"C:\Program Files\wkhtmltopdf\bin\wkhtmltopdf.exe")   

with open(csv_file_path, newline='', encoding="utf8") as csv_file:   
    reader = csv.reader(csv_file)
    header = next(reader)  # 跳过表头
    for row in reader:
        html_code = row[2]
        # 生成唯一文件名
        filename_base = f'{header[2]}_{row[0]}'
        pdf_file_path = f'{pdf_dir_path}\\{filename_base}.pdf'

        try:
            # 直接用HTML字符串转换,无需临时HTML文件
            pdfkit.from_string(html_code, pdf_file_path, configuration=config, options=options)
            print(f'转换成功:{pdf_file_path}')
        
        except Exception as e:
            # 打印具体错误信息,方便排查
            print(f'转换失败:{pdf_file_path},错误信息:{str(e)}')
        

print('所有转换任务完成。')

额外注意事项

  • 确认wkhtmltopdf已正确安装,路径与代码配置完全一致;
  • 若仍有资源加载问题,可尝试添加'no-stop-slow-scripts': None或'disable-javascript': None到options中;
  • 原代码中import pdfkitc为笔误,需改为import pdfkit,否则会触发模块导入错误。

内容的提问来源于stack exchange,提问作者jaypeeee10

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 11:57:52