Python实现Markdown转PDF过程中欧元符号(€)无法正常显示的问题求助
解决Markdown转PDF时欧元符号(€)显示异常的问题
我之前也碰到过类似的字符渲染故障,核心原因就是编码不匹配和wkhtmltopdf的字体/渲染配置缺失,给你几个具体的修复步骤:
1. 给所有文件操作强制指定UTF-8编码
欧元符号属于Unicode字符,默认文件编码大概率无法正确解析它,所以读写Markdown、HTML文件时,必须显式声明UTF-8编码:
修改你的代码片段:
# 写入Markdown文件时 with open(r'path_to_file', 'w', encoding='utf-8') as fp: fp.write(markdown) # 写入HTML文件时 with open('html_format.html', 'w', encoding='utf-8') as f: f.write(html_text)
2. 给生成的HTML添加UTF-8编码声明
markdown库转换出的HTML默认不带编码声明,渲染器可能会用错误的编码解析内容,直接导致欧元符号乱码。在生成的HTML开头补上编码标签:
with open(input_filename, 'r', encoding='utf-8') as f: md_content = f.read() html_content = markdown(md_content, output_format='xhtml', extensions=['markdown.extensions.tables']) # 插入UTF-8编码声明 html_text = f'<meta charset="UTF-8">\n{html_content}'
3. 配置wkhtmltopdf使用支持欧元的字体并指定编码
wkhtmltopdf默认可能使用不支持欧元符号的字体,或者没有正确识别编码,需要在调用时补充配置选项:
config = pdfkit.configuration(wkhtmltopdf=r"C:\Program Files\wkhtmltopdf\bin\wkhtmltopdf.exe") # 定义渲染选项 options = { 'encoding': 'UTF-8', # 指定系统中支持欧元的字体,比如Arial、Helvetica都是系统自带的安全选项 'font-name': 'Arial', 'page-size': 'A4' } # 注意要传入定义好的config和options参数 pdfkit.from_string(html_text, output_filename, configuration=config, options=options, verbose=True)
为什么这些步骤能解决问题?
- UTF-8是唯一能完整覆盖所有Unicode字符(包括€)的编码格式,显式指定后能彻底避免文件读写时的字符损坏;
- HTML的
<meta charset="UTF-8">会告诉渲染器用正确的编码解析内容,避免识别偏差; - 指定支持欧元的系统字体,能确保wkhtmltopdf渲染时找到对应的字符字形,不会显示方框或乱码。
内容的提问来源于stack exchange,提问作者mosfet631
相关产品推荐
相关产品推荐

