You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

将TXT编码从ANSI转WINDOWS-1256时Python代码报错求助

阿拉伯文编码转换问题解决

问题背景

我有一个source.txt文件,内容显示为:

ËÇÈÊÉ
ÇáÊØåíÑ
ÇÓÊåáÇß

对应的正确阿拉伯文是:

ثابتة
التطهير
استهلاك

该文本实际是阿拉伯文,在Notepad++中手动将编码从ANSI改为WINDOWS-1256就能正常显示。因为文件数量较多,我写了一段Python代码批量处理:

with open("source.txt", 'r', encoding='ansi') as file_in:
    text = ""
    for line in file_in:
        text = text+line

with open("ARABIC.txt", 'w', encoding='cp1256') as f:
    f.write(text)

运行代码时出现如下错误:

Traceback (most recent call last):
File "C:\Users\user\Desktop\arabiccode\SHOW_ARABIC.py", line 7, in 
f.write(text)
File "C:\Program Files (x86)\Microsoft Visual Studio\Shared\Python37_64\lib\encodings\cp1256.py", line 19, in encode
return codecs.charmap_encode(input,self.errors,encoding_table)[0]
UnicodeEncodeError: 'charmap' codec can't encode characters in position 0-4: character maps to <undefined>

错误原因

用ansi编码读取文件时,已将原始的cp1256字节错误解码成Unicode字符,这些错误字符在后续用cp1256编码时找不到对应映射,因此抛出编码错误。

修正方案

方案一:直接字节读写(最简单高效)

原始文件的字节本身就是cp1256编码,直接按字节复制即可:

with open("source.txt", 'rb') as file_in:
    raw_bytes = file_in.read()

with open("ARABIC.txt", 'wb') as f:
    f.write(raw_bytes)

方案二:明确编码转换

先以正确的cp1256编码读取文件,再写入:

with open("source.txt", 'r', encoding='cp1256') as file_in:
    text = file_in.read()

with open("ARABIC.txt", 'w', encoding='cp1256') as f:
    f.write(text)

或者通过字节流手动处理解码编码:

with open("source.txt", 'rb') as file_in:
    raw_bytes = file_in.read()
# 用cp1256解码为Unicode文本
text = raw_bytes.decode('cp1256')
# 再用cp1256编码写入文件
with open("ARABIC.txt", 'w', encoding='cp1256') as f:
    f.write(text)

关键提示

  • 你的文件实际编码是WINDOWS-1256(即cp1256),不是系统默认的ANSI编码,用错误编码读取会导致字符乱码,后续无法正常转换。
  • 批量处理时,只需遍历目标文件目录,对每个文件执行上述任一方案即可。

内容的提问来源于stack exchange,提问作者Abid Abdo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 05:07:29