为何int()转换Unicode字符串会出现非预期结果?(如int('᪐᭒')==2)
为何
int('᪐᭒')返回2而非报错 这两个Unicode字符的属性决定了int()的处理结果:
U+1A90(字符᪐)是巴厘语数字2,属于Unicode的Nd(十进制数字)类别,本身对应数值2;U+1B52(字符᭒)是巴厘语数字分隔符,属于No(其他数字)类别,仅用于数字格式分隔,不携带实际数值,会被int()忽略。
Python的int()函数并非只识别阿拉伯数字(0-9),它会解析所有Unicode标准定义的十进制数字字符,同时跳过无数值的数字相关分隔字符。因此传入'᪐᭒'时,函数提取有效数字᪐对应的数值2,最终返回结果。
你可以用unicodedata模块验证字符属性:
import unicodedata # 验证巴厘语数字2 print(unicodedata.name('᪐')) # BALINESE DIGIT TWO print(unicodedata.category('᪐')) # Nd(十进制数字) print(unicodedata.numeric('᪐')) # 2.0 # 验证巴厘语数字分隔符 print(unicodedata.name('᭒')) # BALINESE DIGIT SEPARATOR print(unicodedata.category('᭒')) # No(其他数字) print(unicodedata.numeric('᭒', None)) # None(无对应数值)
如果需要限制仅识别阿拉伯数字,可提前过滤字符串,只保留0-9范围内的字符后再转换。
内容的提问来源于stack exchange,提问作者Wör Du Schnaffzig
相关产品推荐
相关产品推荐

