You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按顺序提取HTML标签内文本并匹配样式后通过PIL绘制带格式文本

带HTML标签文本的PIL绘制实现方案

核心思路

放弃手动遍历兄弟节点的处理方式,改为递归解析HTML节点生成「带样式标记的文本片段」,再基于片段拆分带样式的单词,适配原有自动换行逻辑,PIL本身通过加载不同样式的字体文件实现粗/斜等文本样式。

实现步骤

步骤1:HTML解析为带样式的文本片段

先解码HTML实体,再用BeautifulSoup递归遍历所有节点,自动继承父标签的样式属性,无需手动处理节点前后关系:

import html
from bs4 import BeautifulSoup

def parse_html_to_styled_fragments(html_str):
    # 解码<这类HTML实体为原始标签
    decoded_str = html.unescape(html_str)
    soup = BeautifulSoup(decoded_str, "html.parser")
    fragments = []
    
    def traverse_node(node, current_styles):
        # 匹配标签更新当前样式
        if node.name == "b":
            current_styles["bold"] = True
        elif node.name == "i":
            current_styles["italic"] = True
        # 可扩展支持<u>下划线、<font color>颜色等更多标签
        
        # 处理文本节点
        if isinstance(node, str):
            fragments.append({
                "text": node,
                "styles": current_styles.copy()
            })
        # 递归遍历子节点
        else:
            for child in node.children:
                traverse_node(child, current_styles.copy())
    
    traverse_node(soup, {})
    return fragments

步骤2:准备对应样式的字体文件

PIL无法直接对字体做样式变换,需要提前准备不同样式的同字号字体文件:

from PIL import ImageFont

# 按自己的字体路径修改,可扩展支持多字号
FONT_CONFIG = {
    "normal": ImageFont.truetype("微软雅黑常规.ttc", 16),
    "bold": ImageFont.truetype("微软雅黑粗体.ttc", 16),
    "italic": ImageFont.truetype("微软雅黑斜体.ttc", 16)
}

步骤3:拆分带样式的单词,适配自动换行逻辑

把样式片段拆分为携带样式标记的单词,修改原有绘制逻辑适配样式切换:

def get_styled_words(fragments):
    styled_words = []
    for frag in fragments:
        words = frag["text"].split(" ")
        for idx, word in enumerate(words):
            # 处理连续空格
            if not word and idx != len(words)-1:
                styled_words.append({
                    "text": " ",
                    "styles": frag["styles"]
                })
                continue
            # 保留单词后的空格
            suffix = " " if idx != len(words)-1 else ""
            styled_words.append({
                "text": word + suffix,
                "styles": frag["styles"]
            })
    return styled_words

修改后的自动换行绘制代码:

rect = Rectangle(x,y,width,height)
curx = rect.x
cury = rect.y
# 统一行高,取所有字体的最大高度避免错位
line_height = max([font.getmetrics()[0] for font in FONT_CONFIG.values()])
# 生成带样式的单词列表
styled_words = get_styled_words(parse_html_to_styled_fragments(输入的HTML字符串))

for word_item in styled_words:
    # 根据样式选对应字体
    styles = word_item["styles"]
    if styles.get("bold"):
        font = FONT_CONFIG["bold"]
    elif styles.get("italic"):
        font = FONT_CONFIG["italic"]
    else:
        font = FONT_CONFIG["normal"]
    
    word_width = font.getsize(word_item["text"])[0]
    # 换行判断
    if curx + word_width > rect.x + rect.width:
        cury += line_height
        curx = rect.x
    # 绘制文本
    draw.text((curx, cury), word_item["text"], ImageColor.getcolor(hex, "RGB"), font=font)
    curx += word_width

注意事项

  • 若没有单独的斜体/粗体字体文件,可以对绘制后的文本区域做倾斜、描边处理实现近似效果,精度略低
  • 需要支持颜色、自定义字号的话,只需要在current_styles中添加对应属性,FONT_CONFIG按「字号+样式」组合作为key即可
  • 处理换行符、制表符等特殊字符,可在拆分单词的步骤中增加对应判断逻辑

内容的提问来源于stack exchange,提问作者A.Dumas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 02:15:03