You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C/C++中Shift-JIS多字节字符串的strstr替代方案有哪些?

跨平台实现Shift-JIS字符串的strstr功能方案

方案一:手动实现Shift-JIS专属匹配函数

Shift-JIS的编码规则清晰明确:

  • 单字节字符:范围0x00-0x7F,对应ASCII字符
  • 双字节字符:首字节范围是0x81-0x9F、0xE0-0xEF,第二个字节范围是0x40-0x7E、0x80-0xFC

基于这个规则,你可以自己写一个匹配函数,核心是永远不截断双字节字符,确保匹配时是完整的字符对比,从根源避免标准strstr的误匹配问题。

示例代码如下:

#include <cstdint>
#include <cstring>

// 判断是否为Shift-JIS双字节字符的首字节
bool is_sjis_lead(uint8_t c) {
    return (c >= 0x81 && c <= 0x9F) || (c >= 0xE0 && c <= 0xEF);
}

// 获取当前Shift-JIS字符的字节长度
int sjis_char_len(uint8_t c) {
    return is_sjis_lead(c) ? 2 : 1;
}

const char* sjis_strstr(const char* haystack, const char* needle) {
    if (!haystack || !needle || *needle == '\0') {
        return haystack; // 边界情况处理
    }

    const uint8_t* hs = reinterpret_cast<const uint8_t*>(haystack);
    const uint8_t* nd = reinterpret_cast<const uint8_t*>(needle);

    // 先缓存子串第一个字符的长度和内容
    int nd_first_len = sjis_char_len(*nd);
    uint8_t nd_first[2];
    memcpy(nd_first, nd, nd_first_len);

    while (*hs != '\0') {
        int hs_char_len = sjis_char_len(*hs);
        // 匹配子串首字符
        if (hs_char_len == nd_first_len && memcmp(hs, nd_first, nd_first_len) == 0) {
            const uint8_t* hs_ptr = hs + hs_char_len;
            const uint8_t* nd_ptr = nd + nd_first_len;
            bool match_success = true;

            // 逐字符匹配后续内容
            while (*nd_ptr != '\0') {
                if (*hs_ptr == '\0') {
                    match_success = false;
                    break;
                }
                int h_len = sjis_char_len(*hs_ptr);
                int n_len = sjis_char_len(*nd_ptr);
                if (h_len != n_len || memcmp(hs_ptr, nd_ptr, h_len) != 0) {
                    match_success = false;
                    break;
                }
                hs_ptr += h_len;
                nd_ptr += n_len;
            }

            if (match_success) {
                return reinterpret_cast<const char*>(hs);
            }
        }
        hs += hs_char_len;
    }
    return nullptr;
}

这个函数仅依赖C++标准库,跨平台兼容性拉满,完全不需要第三方依赖。

方案二:使用ICU库

ICU是专注于国际化、多语言编码处理的跨平台库,原生支持CMake集成,适合需要处理多种编码的场景。

大致步骤:

  1. 在CMake配置中引入ICU:find_package(ICU REQUIRED COMPONENTS uc i18n)
  2. 将Shift-JIS格式的const char*字符串转换为ICU的UnicodeString对象
  3. 调用UnicodeString::find()方法执行查找
  4. 若匹配成功,计算对应Shift-JIS字符串的字节偏移,返回原字符串指针

这种方案不用自己维护编码规则,ICU会帮你处理所有编码细节。

方案三:使用Boost.Locale

Boost.Locale是Boost生态中的国际化库,同样支持CMake集成,API设计友好,若你的项目已经在使用Boost,这个方案几乎没有额外接入成本。

大致步骤:

  1. 在CMake中引入Boost.Locale:find_package(Boost REQUIRED COMPONENTS locale)
  2. 创建指定Shift-JIS编码的boost::locale::string对象
  3. 使用boost::algorithm::find()或字符串自身的查找方法执行匹配
  4. 将匹配结果转换回原const char*指针位置

内容的提问来源于stack exchange,提问作者Bowen Cui

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 04:22:53