PHP解析希伯来语字符异常:完整单词正常,拆分后乱码
希伯来语字符拆分显示异常问题解决
问题现象
在PHP中解析希伯来语单词时,完整字符串显示正常,但用substr拆分单个字符后出现乱码(显示为�)。生产环境中单词从采用UTF16_general_ci编码的MySQL数据库获取,测试时使用模拟数据也出现相同问题。添加Content-Type: text/html; charset=UTF-16响应头后,整个页面会显示为乱码(类似简体中文)。
测试代码
<!DOCTYPE html> <?php //uncommenting the next line results in the whole page displaying in "chinese -simplified" //header("content-type: text/html; charset=UTF-16"); header('Content-language: he'); ?> <html> <head> <meta http-equiv=Content-Type content="text/html; charset=UTF-16"> <meta http-equiv="content-language" content="he-il"> </head> <body> <?php // in Production, we are grabbing the hebrew word from the database //$sql = "SELECT masoretic FROM codex WHERE id = 20"; // just grabs a word from the database // it is stored using UTF16_general_ci on mySQL // in this test we can mock the exact same data that was copy and pasted in // the results were the same with the data from the db $masoretic = "בָּרָ֣א"; echo $masoretic . '<br>'; // displays correctly in HEBREW = בָּרָ֣א // now loop through the word and process each letter $length = strlen($masoretic); // even though there are only 3 real letters, the diacritic marks count as characters, so we should get at least 7 loops for ($x = 0; $x <= $length; $x++) { $letter = substr($masoretic,0,1); // process this letter $masoretic = substr($masoretic, 1); // the rest of the word $name = ''; $recognized = false; switch($letter){ case 'ר': $recognized = true; $name = 'Raysh'; break; case 'א': $recognized = true; $name = 'Aleph'; break; default: $recognized = false; break; } if($recognized){ echo ('found a ' . $name); echo $letter; // for now just display it }else{ echo 'unrecognized letter:'; print_r($letter); echo '<br>'; } } ?> </body>
页面输出
בָּרָ֣א unrecognized letter:� unrecognized letter:� unrecognized letter:� unrecognized letter:� unrecognized letter:� unrecognized letter:� unrecognized letter:� unrecognized letter:� unrecognized letter:� unrecognized letter:� unrecognized letter:� unrecognized letter:� unrecognized letter:� unrecognized letter:� unrecognized letter:
问题根源
- 字节级操作不兼容多字节编码:PHP原生的
strlen、substr是基于字节的函数,而UTF-16编码中每个字符占用2或4个字节。拆分时会将单个字符的字节强行拆分,生成无效字节序列,最终显示为乱码�。 - UTF-16兼容性差:浏览器对UTF-16的支持远不如UTF-8广泛,直接设置UTF-16编码容易触发解析错误,导致整页乱码。
- 编码处理不统一:数据库存储为UTF-16,但PHP未做编码转换就直接处理,进一步放大字符拆分的问题。
解决方案
- 使用多字节字符串函数:替换
strlen为mb_strlen、substr为mb_substr,明确指定编码(推荐转成UTF-8处理,兼容性更好)。 - 统一页面编码为UTF-8:将页面响应头和meta标签的编码都设置为UTF-8,避免编码解析冲突。
- 数据库编码适配:从UTF-16编码的数据库读取数据时,先转换为UTF-8再进行处理,确保编码统一。
修正后的代码
<!DOCTYPE html> <?php // 统一设置页面编码为UTF-8 header("content-type: text/html; charset=UTF-8"); header('Content-language: he'); ?> <html> <head> <meta http-equiv=Content-Type content="text/html; charset=UTF-8"> <meta http-equiv="content-language" content="he-il"> </head> <body> <?php $masoretic = "בָּרָ֣א"; // 若从UTF-16数据库读取数据,先转成UTF-8 // $masoretic = mb_convert_encoding($masoretic, 'UTF-8', 'UTF-16'); echo $masoretic . '<br>'; // 使用mb_strlen获取字符长度(而非字节长度) $length = mb_strlen($masoretic, 'UTF-8'); for ($x = 0; $x < $length; $x++) { // 使用mb_substr按字符拆分 $letter = mb_substr($masoretic, $x, 1, 'UTF-8'); $name = ''; $recognized = false; switch($letter){ case 'ר': $recognized = true; $name = 'Raysh'; break; case 'א': $recognized = true; $name = 'Aleph'; break; case 'ב': $recognized = true; $name = 'Bet'; break; // 可继续添加其他希伯来字符的匹配规则 default: $recognized = false; break; } if($recognized){ echo 'found a ' . $name . ': ' . $letter . '<br>'; }else{ echo 'unrecognized character: ' . $letter . '<br>'; } } ?> </body>
内容的提问来源于stack exchange,提问作者ken
相关产品推荐
相关产品推荐

