You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

生成含随机加粗文字的RTL语言图片时水平对齐异常求助

问题描述

我需要生成包含特定文本的图片,用单一字体生成没问题,但想随机加粗部分单词,所以要用到常规和粗体两种字体。我的思路是:

  • 将文本拆分为单词;
  • 随机为单词生成常规或粗体的字体掩码;
  • 因为是从右到左(RTL)语言,要从右到左逐个拼接掩码到同一张图片中。

问题:单词无法保持同一水平对齐!组合不同掩码到单张图片时出现异常。

  • 当前效果:文字上下错位,无法对齐在同一基线
  • 预期效果:所有单词保持水平基线对齐,仅部分单词加粗

实现代码:

from PIL import Image, ImageFont
import random

font_size = 60
font_bold_file = "fonts/Cairo/static/Cairo-Bold.ttf"
font_regular_file = "fonts/Cairo/Cairo-VariableFont_slnt,wght.ttf"
text_color = (0, 0, 0, 255)

font_regular = ImageFont.truetype(font_regular_file, size=font_size)
font_bold = ImageFont.truetype(font_bold_file, size=font_size)

text = "أبت أ ري عالم"
total_x = 0
total_y = 0
segments = []
for w in text.split(" "):

    mask = None
    if random.random() < 1:
        mask = font_regular.getmask(w + " ", "L")
    else:
        mask = font_regular.getmask(w + " ", "L")

    segments.append( (mask, mask.size[0]) )#
    total_x += mask.size[0]
    total_y = max(total_y, mask.size[1])

background_img = Image.new(mode="RGBA", size=(total_x, total_y), color=(255, 255, 0, 255))

acc_widths = 0
for seg in segments:
    mask = seg[0]
    width = seg[1]
    pivot_x = total_x - acc_widths - width
    acc_widths += width

    mask_req = (pivot_x, 0,  pivot_x + width, mask.size[1] )
    background_img.im.paste(text_color, mask_req, mask)


background_img.save("pic.png")
解决方案

你的代码主要有两个核心问题:随机逻辑没生效,以及忽略了字体基线对齐的问题。以下是修正方案:

问题点分析

  1. 随机字体逻辑失效:代码里if random.random() < 1永远为真,根本没用到粗体字体,等于一直用常规字体生成掩码。
  2. 对齐错误根源:直接从顶部粘贴文本块,但常规和粗体字体的基线位置不同(粗体通常上下高度略大),只按顶部对齐会导致错位,必须基于字体的基线来计算Y轴偏移。
  3. RTL文本参数错误:生成掩码时用"L"(左到右)不符合阿拉伯语等RTL语言的排版逻辑,应该用"R"参数。

修正后的代码

from PIL import Image, ImageFont
import random

font_size = 60
font_bold_file = "fonts/Cairo/static/Cairo-Bold.ttf"
font_regular_file = "fonts/Cairo/Cairo-VariableFont_slnt,wght.ttf"
text_color = (0, 0, 0, 255)

font_regular = ImageFont.truetype(font_regular_file, size=font_size)
font_bold = ImageFont.truetype(font_bold_file, size=font_size)

text = "أبت أ ري عالم"
total_x = 0
max_ascent = 0
max_descent = 0
segments = []

for w in text.split(" "):
    # 50%概率随机选择粗体或常规字体
    use_bold = random.random() < 0.5
    current_font = font_bold if use_bold else font_regular
    # 生成RTL文本的掩码(带空格)
    mask = current_font.getmask(w + " ", "R")
    # 获取字体的度量参数:ascent是基线到文字顶部的距离,descent是基线到文字底部的距离
    ascent, descent = current_font.getmetrics()
    # 记录全局最大的ascent和descent,确保画布能容纳所有字体
    max_ascent = max(max_ascent, ascent)
    max_descent = max(max_descent, descent)
    # 保存每个片段的掩码、宽度和当前字体的ascent(用于计算Y偏移)
    segments.append((mask, mask.size[0], ascent))
    total_x += mask.size[0]

# 画布高度 = 最大ascent + 最大descent,覆盖所有字体的上下范围
total_y = max_ascent + max_descent
background_img = Image.new(mode="RGBA", size=(total_x, total_y), color=(255, 255, 0, 255))

acc_widths = 0
for seg in segments:
    mask, width, ascent = seg
    # 从右往左计算X位置,实现RTL排版
    pivot_x = total_x - acc_widths - width
    acc_widths += width
    # 计算Y偏移:用全局最大ascent减去当前字体的ascent,确保所有文字基线对齐
    pivot_y = max_ascent - ascent
    # 定义粘贴区域
    mask_req = (pivot_x, pivot_y, pivot_x + width, pivot_y + mask.size[1])
    background_img.im.paste(text_color, mask_req, mask)

background_img.save("pic.png")

关键修正说明

  • 修复随机字体选择逻辑,真正实现部分单词随机加粗
  • 通过getmetrics()获取字体的基线参数,计算每个文本块的Y轴偏移,确保所有单词对齐到同一基线
  • 调整掩码生成参数为"R",适配RTL语言排版
  • 基于全局最大的字体上下高度计算画布尺寸,避免文字被截断

内容的提问来源于stack exchange,提问作者Taha Magdy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 13:50:13