Python中如何将字符串(含不可见字符)转换为Unicode码点表示
Python 任意字符转Unicode码点转义实现方案
基础实现(完全匹配需求示例格式)
直接遍历字符串每个字符,获取码点后格式化为\u+4位十六进制的形式拼接即可,兼容可见字符、零宽字符等所有特殊字符:
def str_to_codepoint_escape(input_str: str) -> str: return ''.join(f'\\u{ord(char):04x}' for char in input_str)
测试示例
test_content = "abc" print(str_to_codepoint_escape(test_content)) # 输出结果:\u0061\u0062\u0063
扩展说明
- 若需要输出大写十六进制格式,把格式化占位符的
x替换为X即可,例如零宽连字符会输出为\u200D - 若需要处理码点大于0xFFFF的增补平面字符(比如Emoji、生僻古籍汉字),可直接用Python内置的
unicode-escape编码实现:
def full_str_to_codepoint_escape(input_str: str) -> str: return input_str.encode("unicode-escape").decode("ascii")
增补字符测试示例
test_content = "😊" print(full_str_to_codepoint_escape(test_content)) # 输出结果:\U0001f60a
内容的提问来源于stack exchange,提问作者DANIEL1475
相关产品推荐
相关产品推荐

