PyPDF2调用addJS添加中文注释时出现解码乱码问题求助
问题根因
PyPDF2的addJS方法默认将传入的字符串以Latin-1编码写入PDF文件,而PDF内置的JavaScript引擎对双字节中文的识别要求使用带BOM的UTF-16BE编码,直接传入UTF-8格式的中文会被阅读器按Latin-1解析,就会出现你观察到的佀好类乱码。
可行解决方案
方案1:使用JS Unicode转义序列替换中文
直接将JS代码中的中文替换为Unicode转义字符,完全规避编码冲突问题,改造成本最低。
修改后代码如下:
from PyPDF2 import PdfFileWriter, PdfFileReader def Test(): inputPDF = PdfFileReader('./demo/TESTPDFANNOTATION.pdf', "rb") outputPDF = PdfFileWriter() pages = inputPDF.getNumPages() for p in range(pages): outputPDF.addPage(inputPDF.getPage(p)) outputStream = open('./demo/TESTPDFANNOTATIONOUT.pdf', "wb") # 中文"你好"对应Unicode转义为\u4f60\u597d outputPDF.addJS("var annot = this.addAnnot({ \r \ page: 0, \r \ type: 'FreeText', \r \ contents: '\\u4f60\\u597d', \r \ textFont: 'csongl', \r \ textSize: 10, \r \ rect: [200, 300, 200+150, 300+3*12], \r \ width: 1, \r \ alignment: 1 \r \ });") outputPDF.write(outputStream) outputStream.close() return "ok"
如果需要批量转换中文,可以用Python的str.encode('unicode-escape').decode()方法自动生成转义序列。
方案2:直接使用PyPDF2原生注释接口(更稳定)
不通过JS注入的方式添加注释,直接调用PyPDF2原生的注释添加接口,编码可控性更高,代码示例如下:
from PyPDF2 import PdfWriter, PdfReader from PyPDF2.generic import DictionaryObject, NameObject, StringObject, ArrayObject, NumberObject def Test(): reader = PdfReader('./demo/TESTPDFANNOTATION.pdf') writer = PdfWriter() for page in reader.pages: writer.add_page(page) # 构造FreeText注释对象 annot = DictionaryObject() annot.update({ NameObject("/Type"): NameObject("/Annot"), NameObject("/Subtype"): NameObject("/FreeText"), NameObject("/Contents"): StringObject("你好"), NameObject("/RC"): StringObject('<?xml version="1.0"?><body xmlns="http://www.w3.org/1999/xhtml" xmlns:xfa="http://www.xfa.org/schema/xfa-data/1.0/" xfa:contentType="text/html" xfa:APIVersion="Acrobat:11.0.0"><p style="text-align:left; font-family:csongl; font-size:10pt">你好</p></body>'), NameObject("/Rect"): ArrayObject([NumberObject(200), NumberObject(300), NumberObject(350), NumberObject(336)]), NameObject("/F"): NumberObject(4), NameObject("/DA"): StringObject("/csongl 10 Tf 0 g"), NameObject("/Q"): NumberObject(1) }) # 给第一页加注释 writer.pages[0].annotations.append(annot) with open('./demo/TESTPDFANNOTATIONOUT.pdf', "wb") as outputStream: writer.write(outputStream) return "ok"
注意事项
- 首先升级PyPDF2到最新版本,旧版本存在已知的中文编码处理bug,升级命令:
pip install --upgrade PyPDF2 - 确保使用的PDF阅读器已内置或安装了注释指定的
csongl字体,避免字体缺失导致的显示异常
内容的提问来源于stack exchange,提问作者Stanley
相关产品推荐
相关产品推荐

