Django项目用xhtml2pdf生成PDF时阿拉伯语显示黑块的问题
解决xhtml2pdf生成阿拉伯语PDF时字符显示黑块/方框的问题
一、修复xhtml2pdf的字体加载与嵌入问题
1. 确保字体路径正确并强制嵌入
xhtml2pdf对静态文件路径解析存在局限,且默认不嵌入字体,需调整@font-face配置:
@font-face { font-family: RTLFont; /* 使用Django静态文件绝对路径,或本地文件绝对路径 */ src: url('{{ STATIC_ROOT }}/Markazi_Text/static/MarkaziText-Regular.ttf'); embed: true; /* 关键:强制将字体嵌入PDF */ }
2. 优化渲染函数配置
在render_to_pdf中明确编码和默认字体,同时添加错误日志方便排查:
from django.http import HttpResponse from django.template.loader import get_template from xhtml2pdf import pisa from io import BytesIO from django.conf import settings def render_to_pdf(template_src, context_dict={}): context_dict['STATIC_ROOT'] = settings.STATIC_ROOT template = get_template(template_src) html = template.render(context_dict) result = BytesIO() pdf_options = { 'encoding': 'UTF-8', 'default_font': 'RTLFont' } pdf = pisa.pisaDocument(BytesIO(html.encode("utf-8")), result, **pdf_options) if not pdf.err: return HttpResponse(result.getvalue(), content_type='application/pdf') print(pdf.err) # 打印错误信息用于排查 return HttpResponse('生成PDF失败')
3. 验证字体文件完整性
确保使用的字体文件未损坏,可更换Amiri、Noto Naskh Arabic等官方阿拉伯语字体测试。
二、替代方案:使用WeasyPrint(推荐)
WeasyPrint对RTL语言、现代CSS和字体嵌入的支持远优于xhtml2pdf,适合复杂排版场景:
1. 安装依赖
pip install weasyprint markdown
2. 重写PDF渲染函数
from django.http import HttpResponse from django.template.loader import get_template from weasyprint import HTML from django.conf import settings import markdown def render_to_pdf(template_src, context_dict={}): context_dict['STATIC_URL'] = settings.STATIC_URL # 处理Markdown转HTML if 'markdown_content' in context_dict: context_dict['markdown_html'] = markdown.markdown(context_dict['markdown_content']) template = get_template(template_src) html = template.render(context_dict) pdf_bytes = HTML(string=html, base_url=settings.STATIC_URL).write_pdf() response = HttpResponse(pdf_bytes, content_type='application/pdf') response['Content-Disposition'] = 'inline; filename="result.pdf"' return response
3. 简化HTML模板
<!DOCTYPE html> <html dir="rtl" lang="ar"> <head> <meta charset="UTF-8"> <style> @font-face { font-family: 'Noto Naskh Arabic'; src: url('{{ STATIC_URL }}/fonts/NotoNaskhArabic-Regular.ttf'); } body { font-family: 'Noto Naskh Arabic', serif; font-size: 16px; } </style> </head> <body> <div>{{ entry.arabic_content }}</div> <!-- 渲染转义后的Markdown内容 --> <div>{{ markdown_html|safe }}</div> </body> </html>
三、替代方案:使用FPDF2
FPDF2专门支持阿拉伯语连字和RTL排版,适合精细控制PDF内容的场景:
1. 安装依赖
pip install fpdf2 arabic-reshaper python-bidi markdown
2. 编写PDF生成函数
from django.http import HttpResponse from fpdf import FPDF import arabic_reshaper from bidi.algorithm import get_display import markdown class ArabicPDF(FPDF): def header(self): self.set_font('NotoNaskhArabic', '', 16) self.cell(0, 10, 'النتائج', 0, 1, 'R') def add_arabic_text(self, text): reshaped_text = arabic_reshaper.reshape(text) bidi_text = get_display(reshaped_text) self.set_font('NotoNaskhArabic', '', 12) self.multi_cell(0, 10, bidi_text) self.ln() def generate_arabic_pdf(request): pdf = ArabicPDF() pdf.add_page() # 添加阿拉伯语字体 pdf.add_font('NotoNaskhArabic', '', f'{settings.STATIC_ROOT}/fonts/NotoNaskhArabic-Regular.ttf', uni=True) theentry = entry.objects.get(id=2) # 添加阿拉伯语内容 pdf.add_arabic_text(theentry.arabic_content) # 添加Markdown转HTML后的内容 markdown_html = markdown.markdown(theentry.markdown_content) pdf.write_html(markdown_html) response = HttpResponse(pdf.output(dest='S').encode('latin-1'), content_type='application/pdf') response['Content-Disposition'] = 'inline; filename="arabic_result.pdf"' return response
内容的提问来源于stack exchange,提问作者Mostafa Mohamed
相关产品推荐
相关产品推荐

