Python生成含德语特殊字符的LuaLaTeX文件报UTF-8无效序列错误
问题场景
用Python编写自动化脚本生成、编译包含ä、ü、ß等德语特殊字符的LuaLaTeX文件时,运行抛出编码错误,错误信息如下:
! String contains an invalid utf-8 sequence.
可复现问题的示例代码:
import subprocess import shutil txtFileRecipe = open(r"C:\Users\canna\OneDrive\Desktop\TestTest.tex", "w") txtFileRecipe.write( ("\\documentclass[a5paper]{article}\n" "\\usepackage[ngerman]{babel}\n" "\\usepackage{fontspec}\n" "\\begin{document}\n" "Äpfelmüß\n" "\\end{document}\n") ) txtFileRecipe.close() subprocess.check_call(["LuaLatex", r"C:\Users\canna\OneDrive\Desktop\TestTest.tex"])
故障原因
Windows环境下Python调用open()写入文件时,默认使用系统本地编码(中文系统下为GBK/CP936)保存文件,而LuaLaTeX默认以UTF-8编码读取.tex源文件。文件实际编码和LuaLaTeX预期编码不匹配,读取到非UTF-8的德语特殊字符字节时就会抛出无效序列错误。
修复方案
写入.tex文件时显式指定encoding='utf-8'参数,确保生成的源文件为标准UTF-8编码即可解决问题。同时推荐使用with上下文管理器处理文件IO,避免文件句柄泄漏;调用LuaLaTeX时可指定cwd参数将编译生成的辅助文件统一输出到tex文件所在目录,避免污染脚本运行目录。
修复后的完整代码:
import subprocess import os tex_path = r"C:\Users\canna\OneDrive\Desktop\TestTest.tex" tex_content = """\\documentclass[a5paper]{article} \\usepackage[ngerman]{babel} \\usepackage{fontspec} \\begin{document} Äpfelmüß \\end{document} """ # 显式指定UTF-8编码写入文件 with open(tex_path, "w", encoding="utf-8") as f: f.write(tex_content) # 调用LuaLaTeX编译,指定工作目录为tex文件所在文件夹 subprocess.check_call( ["lualatex", tex_path], cwd=os.path.dirname(tex_path) )
补充说明:如果使用较旧版本的LuaLaTeX,可在tex文件开头添加% !TEX encoding = UTF-8 Unicode元注释显式声明文件编码,新版本无需该配置。
内容的提问来源于stack exchange,提问作者CRizz
相关产品推荐
相关产品推荐

