Python中UTF-8文件转cp1252时特殊字符处理报错求助
问题分析与解决方案
错误原因
你的代码存在三个核心问题:
- 字节值映射错误:你混淆了Unicode码点和cp1252的字节值,比如
é在cp1252中对应的字节是十进制233(0xE9),而非你写的130(0x82),后者是cp1252的不可编码控制字符,这才触发了UnicodeEncodeError。 - 循环逻辑缺陷:多个
if语句无elif分支,且else会在任意if不满足时执行,导致字符被重复添加到结果中。 - 冗余手动转换:Python原生支持UTF-8到cp1252的编码转换,完全不需要手动映射法语特殊字符——这些字符都是cp1252标准支持的。
最简解决方案
直接利用Python的编码转换能力,无需手动处理字符,代码如下:
import io src_path = '0189enr.asc' dst_path = 'a-new-file2.txt' # 读取UTF-8文件,得到Unicode字符串 with io.open(src_path, mode="r", encoding="utf8") as fd: content = fd.read() # 直接将Unicode字符串编码为cp1252写入文件 with io.open(dst_path, mode="w", encoding="cp1252") as fd: fd.write(content)
若需手动处理特殊字符(可选)
如果确实需要针对特定字符做自定义映射,要直接操作字节序列,避免Unicode编码冲突,示例代码:
src_path = '0189enr.asc' dst_path = 'a-new-file2.txt' # 正确映射:Unicode字符 → cp1252字节值 char_map = { 'è': 138, # cp1252中è的字节值(十进制) 'ô': 147, # cp1252中ô的字节值(十进制) 'é': 233 # cp1252中é的正确字节值(十进制) } with open(src_path, 'r', encoding='utf8') as fd: content = fd.read() out_bytes = b'' for char in content: if char in char_map: out_bytes += bytes([char_map[char]]) else: # 其他字符自动用cp1252编码 out_bytes += char.encode('cp1252') # 以二进制模式写入字节序列 with open(dst_path, 'wb') as fd: fd.write(out_bytes)
内容的提问来源于stack exchange,提问作者openastronomer
相关产品推荐
相关产品推荐

