如何在Python 2.7中以u{变量}形式打印Unicode字符?
如何通过Unicode码点在控制台显示对应字符?
我懂你遇到的问题了——之前直接打印Unicode字符没问题,但现在要处理存成列表的码点,循环读取后没法直接转成对应的字符显示对吧?
你之前能正常输出字符的写法是这样的:
print u'à' # 或者 a = u'à' print a
但现在你的需求是,从文件里读取一行行的Unicode码点(比如文件里存的是224或者U+00E0),让控制台显示对应的à,而不是打印码点本身。
针对这个需求,分两种常见情况给你解决方案:
情况1:文件里是十进制码点(每行一个数字)
根据Python版本用对应的字符转换函数即可:
# Python 2.x 写法 with open("someFileWithAListOfUnicodeCodePoints") as uniCodeFile: for codePoint in uniCodeFile: # 先清理换行符,转成整数类型 cp = int(codePoint.strip()) print unichr(cp) # unichr() 把十进制码点转成Unicode字符 # Python 3.x 写法 with open("someFileWithAListOfUnicodeCodePoints", encoding="utf-8") as uniCodeFile: for codePoint in uniCodeFile: cp = int(codePoint.strip()) print(chr(cp)) # Python3 里 chr() 直接支持所有Unicode码点
情况2:文件里是十六进制码点(带/不带U+前缀)
如果码点是十六进制格式,需要先处理掉前缀再转成整数:
# Python 2.x 写法 with open("someFileWithAListOfUnicodeCodePoints") as uniCodeFile: for codePoint in uniCodeFile: cp_str = codePoint.strip().upper() # 去掉U+前缀(如果存在的话) if cp_str.startswith('U+'): cp_str = cp_str[2:] # 十六进制字符串转整数,第二个参数传16 cp = int(cp_str, 16) print unichr(cp) # Python 3.x 写法 with open("someFileWithAListOfUnicodeCodePoints", encoding="utf-8") as uniCodeFile: for codePoint in uniCodeFile: cp_str = codePoint.strip().upper() if cp_str.startswith('U+'): cp_str = cp_str[2:] cp = int(cp_str, 16) print(chr(cp))
小提示:解决控制台乱码问题
如果是Python 2.x控制台显示乱码,可以在代码开头加几行设置编码:
import sys reload(sys) sys.setdefaultencoding('utf-8')
要是Windows的CMD里显示不对,可以先输入chcp 65001切换到UTF-8编码,再运行脚本。
内容的提问来源于stack exchange,提问作者Zaid Tariq
相关产品推荐
相关产品推荐

