Groovy不支持字符的编码结果不匹配问题及技术分析
Encoding Mismatch with Em Dash「―」in Groovy & PHP
Alright, let's walk through this encoding mismatch issue you're hitting with the em dash character 「―」—here's a breakdown of what's going on and the core hurdles:
Core Problem
The em dash 「―」 is fully supported in UTF-8, but it's not included in the EUC-JP character set—this is the root cause of your encoding mismatches across languages.
PHP's Encoding Behavior
In PHP, when you convert a string containing this em dash to different encodings, you get distinct byte sequences:
- When converting to EUC-JP (using encoding conversion logic paired with
var_dump(input_string)), the resulting byte array is:[161, 189, 10]Note: The
int(10)at index 3 is a newline character - When converting to UTF-8, the byte array becomes:
[226, 128, 141, 10]Note: The
int(10)at index 4 is a newline character
Key Observations & Groovy Hurdle
- EUC-JP doesn't have a dedicated code point for the em dash, so PHP is likely substituting it with a similar "horizontal bar" character (the byte pair
161, 189maps to this stand-in in EUC-JP) - The problem in Groovy comes down to how it handles unsupported characters during encoding. Unlike PHP which might silently substitute the em dash, Groovy could throw an error, drop the character entirely, or generate an unexpected byte sequence when trying to encode
「―」to EUC-JP—leading to the mismatch you're seeing.
内容的提问来源于stack exchange,提问作者Md. Hasan Basri
相关产品推荐
相关产品推荐

