如何编写兼容Python 2/3的单源代码实现内存字符串写入文本文件
兼容Python 2.7和3.x的字符串写入解决方案
首先纠正一个小误解:Python 2中的str(字节串)其实是有decode()方法的——你大概率是不小心对unicode类型调用了decode()才会报错,毕竟只有字节串需要解码为Unicode,而Unicode字符串本身不需要这一步操作。
针对你的场景,我们可以用一个无版本分支、优雅通用的方法处理所有写入前的字符串类型转换,完美兼容两个Python版本:
1. 定义通用的Unicode转换函数
这个函数会自动识别输入类型:在Python 2中将字节串(str)解码为Unicode,在Python 3中将字节串(bytes)解码为字符串,已经是Unicode/字符串的内容直接返回即可:
def to_unicode(s): if isinstance(s, bytes): return s.decode('utf-8') return s
2. 修改写入代码
把所有unicode()强制转换替换成这个函数调用就搞定了:
import io import json import yaml with io.open(some_path, 'w', encoding='utf-8') as the_file: the_file.write(to_unicode(json.dumps(some_object, indent=2))) with io.open(some_path, 'w', encoding='utf-8') as the_file: the_file.write(to_unicode(yaml.dump(some_object, default_flow_style=False))) with io.open(some_path, 'w', encoding='utf-8') as the_file: the_file.write(to_unicode(some_multiline_string))
3. 额外优化:让字符串字面量默认是Unicode
在文件顶部添加这行导入,可以让Python 2中的字符串字面量(比如你的some_multiline_string)默认是unicode类型,和Python 3保持一致,进一步减少转换需求:
from __future__ import unicode_literals
针对库的可选优化
如果想彻底避免类型转换,还可以直接让库返回Unicode:
- JSON:调用
json.dumps时添加ensure_ascii=False参数,Python 2中会直接返回unicode类型:json.dumps(some_object, indent=2, ensure_ascii=False) - PyYAML:调用
yaml.dump时添加encoding=None参数,Python 2中返回的是unicode而非字节串:yaml.dump(some_object, default_flow_style=False, encoding=None)
不过上面的to_unicode函数是最通用的方案,不管库返回什么类型都能处理,也不需要记住各个库的参数细节。
内容的提问来源于stack exchange,提问作者MrCranky
相关产品推荐
相关产品推荐

