You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python批量解码TXT文件时UTF-8解码报错,如何保留特殊字符?

解决方案

1. 自动检测文件编码(推荐)

不同文件可能采用了不同编码规范,前1600个文件是UTF-8编码,剩余文件大概率是Windows-1252(0x80字节在该编码中对应€符号)或其他编码。可以用chardet库自动检测每个文件的编码后再解码:

  • 先安装chardet:
pip install chardet
  • 修改代码:
import chardet

for index, byte_content in enumerate(corpus[0]):
    # 检测当前文件的编码
    detect_result = chardet.detect(byte_content)
    target_encoding = detect_result['encoding'] or 'utf-8'  # 检测失败时默认用UTF-8
    # 解码并赋值,注意避免链式索引导致的赋值失效
    corpus.iloc[index, 0] = byte_content.decode(target_encoding)

2. 针对0x80字节尝试特定编码

如果检测后发现剩余文件多为Windows-1252编码,也可以直接用try-except分支处理,优先用UTF-8解码,失败则切换到Windows-1252:

for index, byte_content in enumerate(corpus[0]):
    try:
        corpus.iloc[index, 0] = byte_content.decode('utf-8')
    except UnicodeDecodeError:
        # Windows-1252兼容ISO-8859-1,且能正确解析0x80这类特殊字节
        corpus.iloc[index, 0] = byte_content.decode('cp1252')

3. 修复索引赋值的潜在问题

原代码中corpus.iloc[index][0]属于链式索引,可能导致赋值无法正确写入DataFrame,建议统一使用corpus.iloc[index, 0]直接定位单元格完成赋值。

内容的提问来源于stack exchange,提问作者kornpat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 09:46:02