You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取含中日文CSV触发UnicodeEncodeError问题求助

Python读取含中日文CSV时print报错的解决办法

问题场景

尝试读取包含中文(如“三”)和日文(如“さん”)的.csv文件,代码如下:

fname = "data.csv"

rows = []
with open(fname, 'r', encoding="utf-8") as file:
    csvreader = csv.reader(file)
    for row in csvreader:
        rows.append(row)
    print(rows)

尽管读取时指定了utf-8编码,执行时仍报错:

Traceback (most recent call last):
  File "c:\Users\m8\Desktop\programing_stuff\python-stuff\handwritten_kanji_recognition - 30-08-2022\app.py", line 25, in <module>
    print(rows)
  File "C:\Users\m8\AppData\Local\Programs\Python\Python310\lib\encodings\cp1252.py", line 19, in encode
    return codecs.charmap_encode(input,self.errors,encoding_table)[0]
UnicodeEncodeError: 'charmap' codec can't encode character '\u4e00' in position 28: character maps to <undefined>

错误原因

这个错误不是读取文件的编码问题,而是Windows默认控制台编码为cp1252(或GBK),无法识别中日文这类Unicode字符,导致print输出时触发编码失败。

解决办法

  • 方法一:修改print的输出编码或错误处理
    直接指定print的输出编码为utf-8,或者设置错误处理策略:

    import sys
    # 强制用utf-8输出到控制台
    print(rows, file=sys.stdout.buffer, encoding='utf-8')
    # 或者忽略无法编码的字符
    print(rows, errors='ignore')
    # 或者用?替换无法编码的字符
    print(rows, errors='replace')
    
  • 方法二:强制控制台使用utf-8编码
    在代码开头添加以下代码,重定向标准输出的编码:

    import sys
    import io
    sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8')
    
  • 方法三:用IDE运行代码
    VS Code、PyCharm这类IDE的终端默认支持utf-8编码,直接在IDE内运行代码即可避免该问题。

  • 方法四:逐个打印行元素
    直接打印列表可能触发编码问题,换成逐行打印元素:

    for row in rows:
        print(row)
    

内容的提问来源于stack exchange,提问作者somethingidk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 14:50:34