Python中如何将rb模式读取的二进制字符串转为普通字符串?
Hey there, let's work through this UnicodeDecodeError issue you're hitting. Since your file has mixed encodings, strict UTF-8 decoding is choking on those invalid bytes (like the 0x93 you mentioned). Here are a few practical solutions to ditch that b' marker and turn your bytes objects into regular strings without errors:
1. 给decode添加错误处理参数
最简单的办法是告诉Python如何处理无法解码的字节,而不是直接抛出错误。你有两个常用选项:
errors='replace': 把无法解码的字节替换成�字符,这样不会丢失周围的文本errors='ignore': 直接跳过无法解码的字节,可能会丢失少量字符,但其余内容能正常保留
对应代码如下(replace版本):
new_list = [item.decode(encoding='utf-8', errors='replace') for item in new_list]
ignore版本:
new_list = [item.decode(encoding='utf-8', errors='ignore') for item in new_list]
2. 尝试其他常见编码
你遇到的0x93字节是个明显的线索——它在Windows-1252编码里是左双引号字符,这种编码在混合编码或遗留文件里非常常见。先试试这个方案,它可能能完整保留你的原始文本:
new_list = [item.decode(encoding='windows-1252') for item in new_list]
Windows-1252能处理很多UTF-8不兼容的"奇怪"字节,说不定根本不需要错误处理就能正常解码。
3. 构建灵活的多编码 fallback 解码器
如果你不确定哪个编码能适配列表里所有元素,可以写个小函数按顺序尝试多种编码,最后再用错误处理兜底:
def safe_decode(byte_string): # 先尝试UTF-8 try: return byte_string.decode('utf-8') except UnicodeDecodeError: # 失败的话试试Windows-1252 try: return byte_string.decode('windows-1252') except UnicodeDecodeError: # 最后兜底用替换策略 return byte_string.decode('utf-8', errors='replace') # 把函数应用到你的列表上 new_list = [safe_decode(item) for item in new_list]
这种方式能最大程度保留原始文本,同时避免解码报错。
顺便提一句:没有什么"直接去掉b'"的魔法——b'只是Python用来标识字节对象的标记。要把它转成字符串,必须用合适的编码解码。上面的方法正是解决了混合编码导致的解码失败问题。
内容的提问来源于stack exchange,提问作者JChat

