You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP解析希伯来语字符异常:完整单词正常,拆分后乱码

希伯来语字符拆分显示异常问题解决

问题现象

在PHP中解析希伯来语单词时,完整字符串显示正常,但用substr拆分单个字符后出现乱码(显示为�)。生产环境中单词从采用UTF16_general_ci编码的MySQL数据库获取,测试时使用模拟数据也出现相同问题。添加Content-Type: text/html; charset=UTF-16响应头后,整个页面会显示为乱码(类似简体中文)。

测试代码

<!DOCTYPE html>
<?php
    //uncommenting the next line results in the whole page displaying in "chinese -simplified"
    //header("content-type: text/html; charset=UTF-16");
    header('Content-language: he');
?>
<html>
<head>
    <meta http-equiv=Content-Type content="text/html; charset=UTF-16">
    <meta http-equiv="content-language" content="he-il">
</head>
<body>
<?php
        // in Production, we are grabbing the hebrew word from the database
        //$sql = "SELECT masoretic FROM codex WHERE id = 20"; // just grabs a word from the database
                                                            // it is stored using UTF16_general_ci on mySQL
        // in this test we can mock the exact same data that was copy and pasted in
        // the results were the same with the data from the db
            $masoretic = "בָּרָ֣א";

            echo $masoretic . '<br>'; // displays correctly in HEBREW = בָּרָ֣א
            // now loop through the word and process each letter
            $length = strlen($masoretic);
            // even though there are only 3 real letters, the diacritic marks count as characters, so we should get at least 7 loops
            for ($x = 0; $x <= $length; $x++) {
                $letter = substr($masoretic,0,1); // process this letter
                $masoretic = substr($masoretic, 1); // the rest of the word
                $name = '';
                $recognized = false;
                switch($letter){
                    case 'ר':
                        $recognized = true;
                        $name = 'Raysh';
                        break;
                    case 'א':
                        $recognized = true;
                        $name = 'Aleph';
                        break;
                    default:
                        $recognized = false;
                        break;
                }
                if($recognized){
                    echo ('found a ' . $name);
                    echo $letter; // for now just display it
                }else{
                        echo 'unrecognized letter:';
                        print_r($letter);
                        echo '<br>';
                }                       
            }           
?>
</body>

页面输出

בָּרָ֣א
unrecognized letter:�
unrecognized letter:�
unrecognized letter:�
unrecognized letter:�
unrecognized letter:�
unrecognized letter:�
unrecognized letter:�
unrecognized letter:�
unrecognized letter:�
unrecognized letter:�
unrecognized letter:�
unrecognized letter:�
unrecognized letter:�
unrecognized letter:�
unrecognized letter:

问题根源

  1. 字节级操作不兼容多字节编码:PHP原生的strlen、substr是基于字节的函数,而UTF-16编码中每个字符占用2或4个字节。拆分时会将单个字符的字节强行拆分,生成无效字节序列,最终显示为乱码�。
  2. UTF-16兼容性差:浏览器对UTF-16的支持远不如UTF-8广泛,直接设置UTF-16编码容易触发解析错误,导致整页乱码。
  3. 编码处理不统一:数据库存储为UTF-16,但PHP未做编码转换就直接处理,进一步放大字符拆分的问题。

解决方案

  1. 使用多字节字符串函数:替换strlen为mb_strlen、substr为mb_substr,明确指定编码(推荐转成UTF-8处理,兼容性更好)。
  2. 统一页面编码为UTF-8:将页面响应头和meta标签的编码都设置为UTF-8,避免编码解析冲突。
  3. 数据库编码适配:从UTF-16编码的数据库读取数据时,先转换为UTF-8再进行处理,确保编码统一。

修正后的代码

<!DOCTYPE html>
<?php
    // 统一设置页面编码为UTF-8
    header("content-type: text/html; charset=UTF-8");
    header('Content-language: he');
?>
<html>
<head>
    <meta http-equiv=Content-Type content="text/html; charset=UTF-8">
    <meta http-equiv="content-language" content="he-il">
</head>
<body>
<?php
    $masoretic = "בָּרָ֣א";
    // 若从UTF-16数据库读取数据,先转成UTF-8
    // $masoretic = mb_convert_encoding($masoretic, 'UTF-8', 'UTF-16');

    echo $masoretic . '<br>';

    // 使用mb_strlen获取字符长度(而非字节长度)
    $length = mb_strlen($masoretic, 'UTF-8');
    for ($x = 0; $x < $length; $x++) {
        // 使用mb_substr按字符拆分
        $letter = mb_substr($masoretic, $x, 1, 'UTF-8');
        $name = '';
        $recognized = false;
        switch($letter){
            case 'ר':
                $recognized = true;
                $name = 'Raysh';
                break;
            case 'א':
                $recognized = true;
                $name = 'Aleph';
                break;
            case 'ב':
                $recognized = true;
                $name = 'Bet';
                break;
            // 可继续添加其他希伯来字符的匹配规则
            default:
                $recognized = false;
                break;
        }
        if($recognized){
            echo 'found a ' . $name . ': ' . $letter . '<br>';
        }else{
            echo 'unrecognized character: ' . $letter . '<br>';
        }                       
    }           
?>
</body>

内容的提问来源于stack exchange,提问作者ken

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 13:14:52