You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用hb_shape()后Harfbuzz返回错误字形码点问题求助

HarfBuzz hb_shape() 调用后 glyph_info.codepoint 异常问题

现象

调用hb_shape()后,glyph_info[i].codepoint返回值不是原字符的Unicode码点(比如字符'H'的Unicode码点是72,却返回43);但如果不调用hb_shape(),该字段能正确返回对应字符的Unicode码点。

相关代码

1. 从FT_Face创建hb_font_t

FT_Library ft;
FT_Init_FreeType(&ft);
FT_Face face;
FT_New_Face(ft, path.c_str(), 0, &face);
FT_Set_Pixel_Sizes(face, 0, size);

_ft_face = face;
_hb_font = hb_ft_font_create(_ft_face, NULL);

2. 字形布局函数

std::vector<Glyph> Font::Shape(hb_buffer_t* buf)
{
    std::vector<Glyph> temp;
    hb_shape(_hb_font, buf, NULL, 0);

    unsigned int glyph_count;
    hb_glyph_info_t* glyph_info = hb_buffer_get_glyph_infos(buf, &glyph_count);
    hb_glyph_position_t* glyph_pos = hb_buffer_get_glyph_positions(buf, &glyph_count);

    Glyph g;
    hb_position_t cursor_x = 0;
    hb_position_t cursor_y = 0;
    for (unsigned int i = 0; i < glyph_count; ++i) {
        hb_codepoint_t glyphid = glyph_info[i].codepoint;
        hb_position_t x_offset = glyph_pos[i].x_offset >> 6;
        hb_position_t y_offset = glyph_pos[i].y_offset >> 6;
        hb_position_t x_advance = glyph_pos[i].x_advance >> 6;
        hb_position_t y_advance = glyph_pos[i].y_advance >> 6;

        //获取预先生成的字形数据(即字体 atlas 中的位置和大小)
        g = _glyphs->at(glyphid); 
        g.pos = glm::vec2(cursor_x + x_offset, cursor_y + y_offset);
        temp.push_back(g);

        cursor_x += x_advance;
        cursor_y += y_advance;
    }

    return temp;
}

3. 缓冲区创建代码

char* text = "Hello, world!";
hb_buffer_t* buf;
buf = hb_buffer_create();
hb_buffer_add_utf8(buf, text, -1, 0, -1);
hb_buffer_set_direction(buf, HB_DIRECTION_LTR);
hb_buffer_set_script(buf, HB_SCRIPT_LATIN);
hb_buffer_set_language(buf, hb_language_from_string("en", -1));

原因

这是HarfBuzz的正常设计行为:

  • 调用hb_shape()前,glyph_info[i].codepoint存储的是输入字符的Unicode码点;
  • 调用hb_shape()后,该字段会被替换为当前字体文件中的字形索引(Glyph ID)——你看到的43,就是当前字体里对应'H'的字形编号,而非Unicode码点。

解决方案

如果需要保留原Unicode码点,有两种可行方式:

  1. 在调用hb_shape()前,提前复制保存glyph_info中的Unicode码点;
  2. 利用glyph_info[i].cluster字段:HarfBuzz会将原字符的索引关联到cluster,通过这个索引可以回溯原输入字符串的Unicode码点。

以下是第一种方式的修改示例:

std::vector<Glyph> Font::Shape(hb_buffer_t* buf)
{
    std::vector<Glyph> temp;
    // 提前保存原Unicode码点
    unsigned int pre_shape_count;
    hb_glyph_info_t* pre_glyph_info = hb_buffer_get_glyph_infos(buf, &pre_shape_count);
    std::vector<hb_codepoint_t> original_codepoints(pre_shape_count);
    for (unsigned int i = 0; i < pre_shape_count; ++i) {
        original_codepoints[i] = pre_glyph_info[i].codepoint;
    }

    hb_shape(_hb_font, buf, NULL, 0);

    unsigned int glyph_count;
    hb_glyph_info_t* glyph_info = hb_buffer_get_glyph_infos(buf, &glyph_count);
    hb_glyph_position_t* glyph_pos = hb_buffer_get_glyph_positions(buf, &glyph_count);

    Glyph g;
    hb_position_t cursor_x = 0;
    hb_position_t cursor_y = 0;
    for (unsigned int i = 0; i < glyph_count; ++i) {
        hb_codepoint_t glyphid = glyph_info[i].codepoint; // 当前是字形索引,用于获取atlas数据
        hb_codepoint_t original_code = original_codepoints[glyph_info[i].cluster]; // 原Unicode码点
        hb_position_t x_offset = glyph_pos[i].x_offset >> 6;
        hb_position_t y_offset = glyph_pos[i].y_offset >> 6;
        hb_position_t x_advance = glyph_pos[i].x_advance >> 6;
        hb_position_t y_advance = glyph_pos[i].y_advance >> 6;

        g = _glyphs->at(glyphid); 
        g.pos = glm::vec2(cursor_x + x_offset, cursor_y + y_offset);
        // 若需要,可将原Unicode码点存入Glyph结构体
        // g.unicode = original_code;
        temp.push_back(g);

        cursor_x += x_advance;
        cursor_y += y_advance;
    }

    return temp;
}

内容的提问来源于stack exchange,提问作者Cortex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 06:20:43