You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python(ReportLab)中识别TXT文件里的≥、≤符号?

解决ReportLab生成PDF时的符号处理问题

一、让Python正确识别TXT中的≤符号

Python读取TXT文件时,只要指定正确的编码(比如utf-8),就能直接识别≤这类Unicode符号。示例代码:

# 读取TXT文件,指定utf-8编码确保特殊符号被正确识别
with open("your_text.txt", "r", encoding="utf-8") as f:
    content = f.read()

如果你的TXT文件采用其他编码(比如gbk),替换对应的编码参数即可,读取后content变量里的≤符号会被正常保留。

二、将>=替换为≥并在PDF中显示

读取文本后先做字符串替换,再用ReportLab写入PDF。同时要注意字体支持:默认的Helvetica字体可能没有≥、≤这类符号,需要换成支持Unicode的字体(比如系统自带的宋体、微软雅黑,或者开源的Noto Sans)。

完整示例代码

from reportlab.pdfgen import canvas
from reportlab.pdfbase import pdfmetrics
from reportlab.pdfbase.ttfonts import TTFont

# 注册支持特殊符号的字体(以Windows系统宋体为例)
pdfmetrics.registerFont(TTFont("SimSun", "C:/Windows/Fonts/simsun.ttc"))

# 读取TXT内容并替换>=为≥
with open("your_text.txt", "r", encoding="utf-8") as f:
    content = f.read().replace(">=", "≥")

# 创建PDF文档
c = canvas.Canvas("output.pdf")
# 设置字体为注册的宋体,字号12
c.setFont("SimSun", 12)

# 逐行写入文本(简单排版)
lines = content.split("\n")
y_position = 750  # 页面起始Y坐标
for line in lines:
    c.drawString(50, y_position, line)
    y_position -= 15  # 每行间距15
    if y_position < 50:  # 超出页面则新建一页
        c.showPage()
        y_position = 750

c.save()

复杂排版优化

如果需要自动换行、对齐等复杂排版,建议使用ReportLab的Paragraph组件:

from reportlab.platypus import SimpleDocTemplate, Paragraph
from reportlab.lib.styles import getSampleStyleSheet

doc = SimpleDocTemplate("output.pdf")
styles = getSampleStyleSheet()
style = styles["BodyText"]
style.fontName = "SimSun"  # 替换为已注册的字体名

# 把处理后的内容转为Paragraph对象
elements = [Paragraph(content, style)]
doc.build(elements)

系统字体适配

  • Linux/macOS系统需替换字体路径,比如macOS的宋体路径为/Library/Fonts/SimSun.ttf
  • 若找不到系统字体,可下载开源Unicode字体(如Noto Sans),将路径指向本地字体文件即可

内容的提问来源于stack exchange,提问作者Ian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 10:01:43