Django xhtml2pdf生成含阿拉伯语PDF报latin-1编码错误
问题报错
触发错误:UnicodeEncodeError: 'latin-1' codec can't encode characters in position 2082-2084: ordinal not in range(256)
问题现象
- Django项目使用xhtml2pdf库生成PDF时,仅传入英文内容可正常生成文件,添加阿拉伯语内容即触发上述编码错误
- 完全相同的代码在其他项目中可正常运行,当前项目手动将编码改为utf-8后,不再抛出编码错误,但生成的PDF页面中阿拉伯语内容显示为空白方框
根因定位
问题由两个独立配置缺失共同导致:
- 编码逻辑错误:
utils.py中硬编码将渲染完成的HTML按ISO-8859-1(即latin-1)编码,该编码仅支持0-255范围的西文字符,完全不覆盖阿拉伯语对应的Unicode码位,因此直接抛出编码异常。 - 字体缺失:xhtml2pdf渲染PDF时不会自动调用系统字体做字形兜底,手动改完utf-8编码后,由于项目中没有配置支持阿拉伯语字形的字体,无法匹配阿拉伯语字符的渲染规则,因此显示为空白方框。
- 同代码在其他项目可正常运行的原因:其他运行环境中要么提前配置了支持阿拉伯语的字体映射和正确的编码逻辑,要么环境内存在xhtml2pdf可默认加载的阿拉伯语字体,因此不会触发问题。
可行解决方案
按以下步骤逐一调整配置即可修复:
- 修正编码处理逻辑
删除utils.py中手动将HTML转为ISO-8859-1编码的代码,调用pisa.pisaDocument时显式指定utf-8编码,参考修改:
import io from django.http import HttpResponse from django.template.loader import render_to_string from xhtml2pdf import pisa def render_to_pdf(template_path, context={}): html = render_to_string(template_path, context) result = io.BytesIO() # 显式指定utf-8编码,不要手动用latin-1编码HTML pdf = pisa.pisaDocument( io.BytesIO(html.encode("UTF-8")), result, encoding="UTF-8", link_callback=pdf_link_callback ) if not pdf.err: return HttpResponse(result.getvalue(), content_type="application/pdf") return None
- 配置静态资源加载回调
在utils.py中添加静态资源路径解析回调,保证xhtml2pdf能正常加载项目中的字体、图片等静态文件:
import os from django.conf import settings def pdf_link_callback(uri, rel): # 解析静态文件路径 if uri.startswith(settings.STATIC_URL): path = os.path.join(settings.STATIC_ROOT, uri.replace(settings.STATIC_URL, "")) # 兼容本地开发环境未执行collectstatic的场景 if not os.path.exists(path): for static_dir in settings.STATICFILES_DIRS: candidate_path = os.path.join(static_dir, uri.replace(settings.STATIC_URL, "")) if os.path.exists(candidate_path): path = candidate_path break return path return uri
- 引入并注册支持阿拉伯语的字体
- 下载支持阿拉伯语的字体文件(例如Noto Naskh Arabic),将常规、粗体字重的ttf文件放到项目静态目录的
fonts文件夹下,例如static/fonts/NotoNaskhArabic-Regular.ttf、static/fonts/NotoNaskhArabic-Bold.ttf - 在
gift_pdf.html的样式块中注册字体,同时配置阿拉伯语从右到左的书写规则:
@font-face { font-family: 'ArabicSupportFont'; src: url('/static/fonts/NotoNaskhArabic-Regular.ttf') format('truetype'); font-weight: normal; } @font-face { font-family: 'ArabicSupportFont'; src: url('/static/fonts/NotoNaskhArabic-Bold.ttf') format('truetype'); font-weight: bold; } /* 全局指定渲染字体 */ * { font-family: 'ArabicSupportFont', sans-serif; } /* 阿拉伯语文本单独配置RTL排版规则 */ .arabic-content { direction: rtl; unicode-bidi: bidi-override; }
- 模板中所有阿拉伯语静态文本、动态渲染的阿拉伯语字段(例如
receiverNameDonatedDonationPage)都加上arabic-content类名,保证排版顺序正确。
- 验证逻辑
调整完成后重启服务,分别验证纯英文、混合阿拉伯语、纯阿拉伯语内容的PDF生成效果,确认无编码报错、无空白方框、文字阅读顺序正常即可。如果仍有个别字符显示异常,更换字符覆盖度更高的阿拉伯语字体即可。
内容的提问来源于stack exchange,提问作者user9479133
相关产品推荐
相关产品推荐

